Reading view

EFF and ARTICLE 19 Submission to the European Commission on the DSA Trusted Flagger Guidelines

EFF and ARTICLE 19 have submitted joint comments to the European Commission on draft guidelines for the Digital Services Act’s trusted flagger mechanism. Having long advocated for a DSA that protects freedom of expression while preserving intermediary liability protections and the prohibition on general monitoring, we welcome the Commission's effort to provide practical guidance on how the trusted flagger system should operate. 

The DSA’s trusted flagger system can help platforms identify illegal content more efficiently. But if implemented poorly, it could also encourage over-removal of lawful speech, weaken due process, and give government authorities disproportionate influence over online expression. 

We support the Commission's focus on good practices and illustrative examples, rather than legal interpretations that could inadvertently steer platforms toward particular enforcement outcomes—and argue that the guidelines should include stronger safeguards to protect freedom of expression, due process, and the impartiality of the trusted flagger system. 

We also support the Commission's clarification that the DSA itself does not define "illegal content"; that determination must come from applicable national or EU law. Trusted flaggers submit prioritized notice, but platforms remain responsible for determining whether content is actually illegal. Platforms must therefore conduct careful, informed assessments and should not assume that a trusted flagger notice necessarily warrants restricting content. 

Our submission highlights several areas where the guidelines could be strengthened: 

  • Cross-border assessments require caution. Platforms should not rely on a trusted flagger notice to assess legality across Member States, where national legal frameworks may differ. 
  • Systemic risks extend beyond content moderation. The DSA's systemic risk framework should not rely too heavily on individual moderation decisions, but should also consider broader platform design choices, including recommender systems. 
  • Law enforcement authorities should generally not be granted trusted flagger status. They already have statutory powers under Article 9 of the DSA, and combining those powers with trusted flagger status creates a risk that platforms may treat trusted flagger notices as de facto removal orders, undermining due process and the rule of law. 
  • Civil society organizations play an essential role. Civil society organizations help identify illegal content and report human rights abuses, but the guidelines should also recognize that these organizations may face retaliation for their work and should be protected from abusive campaigns that threaten their independence. 
  • Trusted flaggers should complement—not replace—existing partnerships. The new mechanism should not sideline existing trusted partnership programs, including collaborations with civil society organizations that do not or cannot hold trusted flagger status, especially those outside of the EU with valuable regional expertise.  

Read the full submission here:

  •  

Automated Moderation Is Here to Stay—Accountability Must Keep Pace

This post is part 2 in a series about automated content moderation. Read the first post here.

When whistleblower Frances Haugen leaked a set of documents from Meta in 2020, among the revelations was a jarring statistic: The company’s algorithms designed to detect terrorist content incorrectly deleted nonviolent Arabic-language content 77 percent of the time, while failing to detect hate speech under the company’s own policies in many instances. Meta’s own transparency report released later that year demonstrated similar findings. Five years later, researchers in the region report that overzealous moderation remains a problem, while paths to remedy have all but collapsed.

Where these systems are faltering in Arabic, they’re positively failing in less-resourced languages. As a 2025 report from the Center for Democracy and Technology found, labeled datasets in certain languages and dialects such as Maghrebi Arabic and Kiswahili contain inconsistencies, bias, and inaccuracies due to the limited hiring of annotators who actually speak the languages as well as shifts in the languages themselves. An investigation into ChatGPT’s outputs in several low-resource languages demonstrates the depth of problem.

But language disparities are just one of several concerns as automated moderation becomes more widespread. From the systemic suppression of content from Palestine to the repeated misclassification of LGBTQ+ content as adult or explicit material, these varied examples demonstrate the risks of overreliance on automated moderation—and the need for stronger safeguards.

Transparency, Cultural Competence, Appeals

As we discussed in Part 1 of this series, automated systems can process content at a scale that humans never could, potentially enabling better moderation at scale and alleviating the psychological load on ill-paid moderators whose jobs require them to view incredibly disturbing content. But automated systems also reproduce existing biases, struggle to understand context, and often make mistakes that disproportionately affect journalists, activists, artists, and other vulnerable and marginalized communities.

As Rachel Griffin wrote in 2023, “Perfectly accurate moderation is not only technically out of reach but intrinsically impossible.” Despite those intrinsic flaws, there is a great deal companies, policymakers, and civil society can do to help ensure that highly-automated systems operate in ways that respect human rights, minimize predictable harms, and provide meaningful accountability when they fail. If companies are going to continue relying on automation to moderate users’ speech—and there is little reason to believe they won’t—then accountability must evolve alongside these technologies.

That evolution can start with committing to the Santa Clara Principles 2.0. These principles, first outlined in 2020 and re-launched in 2021 after substantial international input, reflect the needs and expectations of the global community and specifically address automation. The first Foundational Principle states:

Companies should ensure that human rights and due process considerations are integrated at all stages of the content moderation process, and should publish information outlining how this integration is made. Companies should only use automated processes to identify or remove content or suspend accounts, whether supplemented by human review or not, when there is sufficiently high confidence in the quality and accuracy of those processes. Companies should also provide users with clear and accessible methods of obtaining support in the event of content and account action. 

Drawing on the Santa Clara Principles 2.0, international human rights standards, and years of research documenting the shortcomings of automated moderation, we propose eight recommendations for policymakers thinking about regulation and companies deploying AI-assisted content moderation systems.

  1. Automated technologies should help, not replace, human moderators. For example, automated systems can help flag and prioritize content for review, while humans can interpret context, handle sensitive cases, and refine system performance.
  2. Companies must be transparent about when and how automation is used in content decisions.
  3. Companies must regularly audit their automated systems for bias, with particular attention to low-resource languages, vulnerable and marginalized communities, and conflict zones.
  4. Users must have the ability to appeal, and to provide context when they believe human or automated moderation decisions have wrongfully removed their content. Appeals should be promptly evaluated and decided by human moderators.
  5. Companies should regularly assess the human rights impact of their moderation decisions, and issue public statements of the results
  6. If they rely on third-party vendors, companies should carefully (and regularly) audit those vendors for compliance with these same principles
  7. Lawmakers should avoid promoting and passing legislation that effectively or explicitly mandates automated moderation systems
  8. Policymakers should also refrain from attempting to dictate platforms technical and design choices to favor or disfavor particular expression.

These recommendations understand that automated content moderation isn’t just a technical problem for clever engineers and product teams to solve. Because content moderation shapes public discourse and fundamental rights, its design and oversight must respond to the concerns of policymakers, civil society, independent researchers, and the communities most affected by these systems.

This is the second post in a 2-part series on automated content moderation. Read the first post here.

  •  

Automated Moderation Is Here to Stay

This blog post is part 1 of a 2-part series. The second part sets out recommendations for companies and policymakers.

Six years ago—one month into a global pandemic—we argued that the automated moderation processes many platforms were rapidly adopting should be highly transparent, easily appealable, and temporary. We warned that "protocols adopted in times of crisis often persist when the crisis is over."

That warning proved prescient. The use of automation and artificial intelligence (AI) to identify, flag, and moderate content has become the new norm—a permanent feature of how platforms govern speech online. In this two part series, we’re take stock of this new norm, and considering what platforms can and should do to ensure that AI serves online expression rather than stifling it.

A brief history of automated content moderation

From spam filtering and keyword blacklists to the hash-matching technologies used to identify child sexual abuse material and terrorist content, automated technologies have been used in commercial content moderation for many years. While these tools have long posed risks to freedom of expression, their use was, for quite some time, relatively limited in scope.

Then, in 2017, a blog post published by Facebook (now Meta) described the company's "fairly recent" use of artificial intelligence to identify, classify, and remove violent extremist content. At the same time, Facebook emphasized caution, noting that it did not want to suggest there was "any easy technical fix."

Just one year later, Mark Zuckerberg appeared before the U.S. Senate's Commerce and Judiciary Committees and disclosed that "99 percent of the ISIS and Al Qaida content" removed by Facebook was flagged by AI "before any human sees it." He also stated that Facebook was "developing A.I. tools that can identify certain classes of bad activity proactively and flag it for our team at Facebook." At the time, we raised concerns about the ethical implications of using AI in this manner.

Then came 2020. The sudden reduction of the human moderation workforce, combined with a dramatic increase in social media use—and with it, a surge in misinformation—created the perfect conditions for platforms to expand their reliance on AI-driven moderation. It quickly became apparent that companies'—and particularly Meta's—approach to moderation during the pandemic represented a backslide in transparency, freedom of expression, and access to remedy. The increased reliance on automation was a significant factor.

The costs and benefits of AI content moderation

We knew in 2020 that the use of AI to moderate content would present problems for online freedom of expression. Today, those problems are well-documented. A 2025 joint declaration by special rapporteurs and representatives of the United Nations (UN), Organization for Security and Co-operation in Europe (OSCE), Organization of American States (OAS), and African Commission on Human and Peoples’ Rights (ACHPR) states:

“The use of AI content moderation can lead to over-removal, discrimination and censorship. Reliance on inherently biased datasets and opaque training processes can amplify pre-existing inequalities, risking homogenisation of expression, and erasure of linguistic and cultural diversity.”

EFF and many of our allies have documented these impacts. For example, our 2019 paper co-authored with Witness and Syrian Archive examined the impact of extremist content regulations—and their implementation through automation and AI—on human rights documentation. A 2020 report from Human Rights Watch highlighted the consequences of these removals, noting: "There is no way of knowing how much potential evidence of serious crimes is disappearing without anyone's knowledge."

The Center for Democracy and Technology's recent series on content moderation in the Global South demonstrates persistent inequities in content moderation of four “low-resource” languages—so-called because the relative scarcity of training data makes it more difficult to develop equitable and accurate AI models for them. 

Content moderation often disproportionately impacts vulnerable and historically marginalized groups, and AI content moderation is no different. GLAAD recognizes the role AI plays in scaling content moderation but notes that “when moderation systems lack nuance, transparency, and human oversight, they can fail to curb harassment and wrongly suppress legitimate LGBTQ content.”

These failures are not incidental. They are a predictable consequence of deploying automated systems to make complex judgments about language, culture, context, and identity at scale.

All of that said, automated content moderation can offer important benefits. The primary one: helping to spare human content moderators who must review content that varies from whimsical to horrific, often for little pay and with devastating mental health consequences. Outsourcing this work to the bots can offer some relief—though it’s worth noting that the humans hired to train the AI models face a similar dynamic.

In addition, AI models could potentially be trained over time to be more precise, accurate, and dynamic, helping to mitigate over-censorship and disinformation. The jury is still out on whether this potential will be realized; what we do know is that new approaches to the persistent problem of over and under-enforcement are desperately needed.

Automated moderation is no longer an experiment

Getting the balance between real costs and potential benefits depends a lot on the details: how automated systems are designed, trained, implemented, and audited.  

Despite advances in the sophistication and scale of automated moderation systems, many of the transparency, accountability, and due process safeguards advocated by civil society, researchers, and human rights experts have yet to be fully realized. At the same time, automated systems have become increasingly central to how platforms enforce their rules and govern online speech.

The question today is not whether companies will use AI to moderate content, but under what conditions they should do so. And now as ever, the answer is not that the public should just trust that platforms’ deployment of increasingly powerful systems will serve, rather than inhibit online expression. In fact, as automated systems become more sophisticated and more deeply embedded in platform governance, the need for transparency and accountability becomes more urgent. 

This is part 1 of a 2-part series. You can read the second part here.

  •  

The UK’s New Under-16 Social Media Ban Will Cause More Harm Than It Prevents

This week, politicians in the UK pushed forward with plans to eviscerate privacy and free speech on the internet by announcing a ban on social media for users under 16 that is set to take effect in Spring 2027. 

The UK government continues to falsely characterize this policy as a necessary response to growing concerns about online harms for young people. In reality, much like the Online Safety Act, it will cause more harm than it will prevent. 

Users of all ages are burdened with proving their age before accessing content, with social media platforms such as Snapchat, TikTok, YouTube, Instagram, Facebook, and X included in the ban. There remains no reliable, privacy-preserving method of verifying the age of every internet user and methods vary from one platform to the next.

Young people will not simply be protected from being contacted by adults or endlessly scrolling—they’ll also lose access to educational videos on YouTube, local events on Facebook, and potentially cut off from distant friends and family. 

Public policy must be effective, proportionate and respectful of fundamental rights. Young people deserve better than a policy built on panic, and all internet users deserve a safe and free internet. A social media ban generates headlines, but it will not solve the problem. 

A Brief History of Age-Gating in the UK

Age restriction proposals in the UK date back to a decade ago, when the proposed Digital Economy Bill was put forth to (among other things) restrict young people from accessing pornographic websites. While the Digital Economy Act of 2017 passed without age-based restrictions, it laid the groundwork for later age verification measures.

Over the next few years, age checks for porn websites were announced then delayed several times. But it wasn’t until a consultation under the 2016-2019 May government and the 2020 publication of the Online Harms Whitepaper that age verification became a broader idea.

In 2023, the UK passed the controversial Online Safety Act, establishing powers that could weaken privacy protections and freedom of expression for internet users worldwide. In July 2025, the government implemented age assurance measures on sites hosting “harmful” content. 

And despite politicians affirming repeatedly that the Online Safety Act would solve all of the problems with online safety, this year they decided it in fact did not go far enough. American social psychologist and The Anxious Generation author Jonathan Haidt—who has called for age-related social media bans around the world, despite significant scientific doubt about his research—met with the UK Health Secretary in February to push for the ban.

In March, politicians introduced plans for a social media ban into the Children’s Wellbeing and Schools Bill to “prevent children under the age of 16 from becoming or being users” of “all regulated user-to-user services,” to be implemented by “highly-effective age assurance measures”—effectively banning under-16s from social media. 

When this proposal came before the House of Commons, MPs defeated and proposed their own amendment: enabling the Secretary of State to introduce provisions “requiring providers of specified internet services” to prevent access by children, under age 18 rather than 16, to specified internet services or to specified features; and to restrict access by children to specified internet services which ministers provide. 

But the social media ban does not stop there. The provision also requires internet service providers to limit the time kids spend online, and has rules about who can contact them online. These extreme rules will take decisions about using technology away from families and put them in the hands of government regulators. 

The history of this proposal shows that the UK government has repeatedly returned to the same flawed idea: restricting access to online services by requiring age checks for everyone. But the fundamental problems have not changed. There is still no widely available way to verify age online without compromising privacy—but even if there were, broad restrictions on social media will inevitably limit access to lawful speech, and valuable online communities, and arts and culture.

  •  
❌