Normal view

Automated Moderation Is Here to Stay—Accountability Must Keep Pace

10 July 2026 at 15:19

This post is part 2 in a series about automated content moderation. Read the first post here.

When whistleblower Frances Haugen leaked a set of documents from Meta in 2020, among the revelations was a jarring statistic: The company’s algorithms designed to detect terrorist content incorrectly deleted nonviolent Arabic-language content 77 percent of the time, while failing to detect hate speech under the company’s own policies in many instances. Meta’s own transparency report released later that year demonstrated similar findings. Five years later, researchers in the region report that overzealous moderation remains a problem, while paths to remedy have all but collapsed.

Where these systems are faltering in Arabic, they’re positively failing in less-resourced languages. As a 2025 report from the Center for Democracy and Technology found, labeled datasets in certain languages and dialects such as Maghrebi Arabic and Kiswahili contain inconsistencies, bias, and inaccuracies due to the limited hiring of annotators who actually speak the languages as well as shifts in the languages themselves. An investigation into ChatGPT’s outputs in several low-resource languages demonstrates the depth of problem.

But language disparities are just one of several concerns as automated moderation becomes more widespread. From the systemic suppression of content from Palestine to the repeated misclassification of LGBTQ+ content as adult or explicit material, these varied examples demonstrate the risks of overreliance on automated moderation—and the need for stronger safeguards.

Transparency, Cultural Competence, Appeals

As we discussed in Part 1 of this series, automated systems can process content at a scale that humans never could, potentially enabling better moderation at scale and alleviating the psychological load on ill-paid moderators whose jobs require them to view incredibly disturbing content. But automated systems also reproduce existing biases, struggle to understand context, and often make mistakes that disproportionately affect journalists, activists, artists, and other vulnerable and marginalized communities.

As Rachel Griffin wrote in 2023, “Perfectly accurate moderation is not only technically out of reach but intrinsically impossible.” Despite those intrinsic flaws, there is a great deal companies, policymakers, and civil society can do to help ensure that highly-automated systems operate in ways that respect human rights, minimize predictable harms, and provide meaningful accountability when they fail. If companies are going to continue relying on automation to moderate users’ speech—and there is little reason to believe they won’t—then accountability must evolve alongside these technologies.

That evolution can start with committing to the Santa Clara Principles 2.0. These principles, first outlined in 2020 and re-launched in 2021 after substantial international input, reflect the needs and expectations of the global community and specifically address automation. The first Foundational Principle states:

Companies should ensure that human rights and due process considerations are integrated at all stages of the content moderation process, and should publish information outlining how this integration is made. Companies should only use automated processes to identify or remove content or suspend accounts, whether supplemented by human review or not, when there is sufficiently high confidence in the quality and accuracy of those processes. Companies should also provide users with clear and accessible methods of obtaining support in the event of content and account action. 

Drawing on the Santa Clara Principles 2.0, international human rights standards, and years of research documenting the shortcomings of automated moderation, we propose eight recommendations for policymakers thinking about regulation and companies deploying AI-assisted content moderation systems.

  1. Automated technologies should help, not replace, human moderators. For example, automated systems can help flag and prioritize content for review, while humans can interpret context, handle sensitive cases, and refine system performance.
  2. Companies must be transparent about when and how automation is used in content decisions.
  3. Companies must regularly audit their automated systems for bias, with particular attention to low-resource languages, vulnerable and marginalized communities, and conflict zones.
  4. Users must have the ability to appeal, and to provide context when they believe human or automated moderation decisions have wrongfully removed their content. Appeals should be promptly evaluated and decided by human moderators.
  5. Companies should regularly assess the human rights impact of their moderation decisions, and issue public statements of the results
  6. If they rely on third-party vendors, companies should carefully (and regularly) audit those vendors for compliance with these same principles
  7. Lawmakers should avoid promoting and passing legislation that effectively or explicitly mandates automated moderation systems
  8. Policymakers should also refrain from attempting to dictate platforms technical and design choices to favor or disfavor particular expression.

These recommendations understand that automated content moderation isn’t just a technical problem for clever engineers and product teams to solve. Because content moderation shapes public discourse and fundamental rights, its design and oversight must respond to the concerns of policymakers, civil society, independent researchers, and the communities most affected by these systems.

This is the second post in a 2-part series on automated content moderation. Read the first post here.

Automated Moderation Is Here to Stay

7 July 2026 at 18:21

This blog post is part 1 of a 2-part series. The second part sets out recommendations for companies and policymakers.

Six years ago—one month into a global pandemic—we argued that the automated moderation processes many platforms were rapidly adopting should be highly transparent, easily appealable, and temporary. We warned that "protocols adopted in times of crisis often persist when the crisis is over."

That warning proved prescient. The use of automation and artificial intelligence (AI) to identify, flag, and moderate content has become the new norm—a permanent feature of how platforms govern speech online. In this two part series, we’re take stock of this new norm, and considering what platforms can and should do to ensure that AI serves online expression rather than stifling it.

A brief history of automated content moderation

From spam filtering and keyword blacklists to the hash-matching technologies used to identify child sexual abuse material and terrorist content, automated technologies have been used in commercial content moderation for many years. While these tools have long posed risks to freedom of expression, their use was, for quite some time, relatively limited in scope.

Then, in 2017, a blog post published by Facebook (now Meta) described the company's "fairly recent" use of artificial intelligence to identify, classify, and remove violent extremist content. At the same time, Facebook emphasized caution, noting that it did not want to suggest there was "any easy technical fix."

Just one year later, Mark Zuckerberg appeared before the U.S. Senate's Commerce and Judiciary Committees and disclosed that "99 percent of the ISIS and Al Qaida content" removed by Facebook was flagged by AI "before any human sees it." He also stated that Facebook was "developing A.I. tools that can identify certain classes of bad activity proactively and flag it for our team at Facebook." At the time, we raised concerns about the ethical implications of using AI in this manner.

Then came 2020. The sudden reduction of the human moderation workforce, combined with a dramatic increase in social media use—and with it, a surge in misinformation—created the perfect conditions for platforms to expand their reliance on AI-driven moderation. It quickly became apparent that companies'—and particularly Meta's—approach to moderation during the pandemic represented a backslide in transparency, freedom of expression, and access to remedy. The increased reliance on automation was a significant factor.

The costs and benefits of AI content moderation

We knew in 2020 that the use of AI to moderate content would present problems for online freedom of expression. Today, those problems are well-documented. A 2025 joint declaration by special rapporteurs and representatives of the United Nations (UN), Organization for Security and Co-operation in Europe (OSCE), Organization of American States (OAS), and African Commission on Human and Peoples’ Rights (ACHPR) states:

“The use of AI content moderation can lead to over-removal, discrimination and censorship. Reliance on inherently biased datasets and opaque training processes can amplify pre-existing inequalities, risking homogenisation of expression, and erasure of linguistic and cultural diversity.”

EFF and many of our allies have documented these impacts. For example, our 2019 paper co-authored with Witness and Syrian Archive examined the impact of extremist content regulations—and their implementation through automation and AI—on human rights documentation. A 2020 report from Human Rights Watch highlighted the consequences of these removals, noting: "There is no way of knowing how much potential evidence of serious crimes is disappearing without anyone's knowledge."

The Center for Democracy and Technology's recent series on content moderation in the Global South demonstrates persistent inequities in content moderation of four “low-resource” languages—so-called because the relative scarcity of training data makes it more difficult to develop equitable and accurate AI models for them. 

Content moderation often disproportionately impacts vulnerable and historically marginalized groups, and AI content moderation is no different. GLAAD recognizes the role AI plays in scaling content moderation but notes that “when moderation systems lack nuance, transparency, and human oversight, they can fail to curb harassment and wrongly suppress legitimate LGBTQ content.”

These failures are not incidental. They are a predictable consequence of deploying automated systems to make complex judgments about language, culture, context, and identity at scale.

All of that said, automated content moderation can offer important benefits. The primary one: helping to spare human content moderators who must review content that varies from whimsical to horrific, often for little pay and with devastating mental health consequences. Outsourcing this work to the bots can offer some relief—though it’s worth noting that the humans hired to train the AI models face a similar dynamic.

In addition, AI models could potentially be trained over time to be more precise, accurate, and dynamic, helping to mitigate over-censorship and disinformation. The jury is still out on whether this potential will be realized; what we do know is that new approaches to the persistent problem of over and under-enforcement are desperately needed.

Automated moderation is no longer an experiment

Getting the balance between real costs and potential benefits depends a lot on the details: how automated systems are designed, trained, implemented, and audited.  

Despite advances in the sophistication and scale of automated moderation systems, many of the transparency, accountability, and due process safeguards advocated by civil society, researchers, and human rights experts have yet to be fully realized. At the same time, automated systems have become increasingly central to how platforms enforce their rules and govern online speech.

The question today is not whether companies will use AI to moderate content, but under what conditions they should do so. And now as ever, the answer is not that the public should just trust that platforms’ deployment of increasingly powerful systems will serve, rather than inhibit online expression. In fact, as automated systems become more sophisticated and more deeply embedded in platform governance, the need for transparency and accountability becomes more urgent. 

This is part 1 of a 2-part series. You can read the second part here.

❌