RSS Amplifier

Fairly AI · Aug 16, 2026

AI Watermarking or AI Stigmatizing?

0
Sign in to vote or save

Wei Chen · Fairly AI

Disclaimer: This blog reflects my personal opinion and does not constitute legal advice.

The transparency obligations under Article 50 of the EU AI Act took effect on August 2, 2026. Put simply, the law requires (1) certain AI systems to identify themselves when interacting with humans, and (2) certain AI-generated content to be labeled as such.

With Article 50 taking effect, leading AI companies and publishing platforms rushed to roll out features that make AI-generated content easier to identify. Most notably:

  • Anthropic adds invisible watermarks to text generated by Claude. For image files, it uses cryptographically signed notes.

  • Google and OpenAI apply invisible watermarks on AI-generated images and audio, and in the case of Gemini, on music and videos as well.

  • Substack partnered with Pangram to assign a score (1-100%) to blogs pronouncing how much of the blog was written by AI.

  • LinkedIn adds an option for users to flag a post as “Seems like AI slop“ alongside “Report comment“ or “I don’t want to see this.”

Soon, leading AI companies will offer APIs that allow developers to build watermark detection capabilities in many products. This rush toward AI detection is increasing the stigma associated with AI-generated content. It would not be a stretch to imagine, in the coming months:

  • Customers start blocking marketing emails drafted using AI.

  • Employees stop using AI due to the fear that their legal department would require them to label everything touched by AI.

  • Platforms refuse to surface white papers, case studies or client alerts that are technically insightful but written with the assistance of AI.

I do not think that is the right way to read or apply the law.

The EU AI Act is not launching a war against technical advancement or efficiency. It is launching a war against the misuse of human trust.

The EU AI Act’s focus on misuse of human trust is apparent from the legislative text. Under Article 50, if you are a deployer (aka user) of an AI system, your obligations are limited to a few use cases:

  • Emotional/Biometric Use: Deployers of an emotion recognition system or a biometric categorisation system shall inform the natural persons exposed thereto of the operation of the system, and shall process the personal data in accordance with Regulations (EU) 2016/679 and (EU) 2018/1725 and Directive (EU) 2016/680, as applicable.

  • Deepfake: Deployers of an AI system that generates or manipulates image, audio or video content constituting a deep fake, shall disclose that the content has been artificially generated or manipulated.

  • Public Interest Text: Deployers of an AI system that generates or manipulates text which is published with the purpose of informing the public on matters of public interest shall disclose that the text has been artificially generated or manipulated.

Put the legalese into examples:

In the three examples above, the trust in the output depends on the human behind it. Faking the human presence is a betrayal of trust. The same can be said about creative work, such as thought leadership blogs, training courses, novels or paintings, where readers seek a genuine connection with an author’s distinct voice, lived experience, and personal perspective.

The need for human presence, on the other hand, is not the same for a press release, a dashboard showing business performance, a registration statement or a contract. Article 50 does not require a deployer of an AI system to label such content. Most people likely would not be offended to know that AI was used so long as the content is accurate, substantive and relevant. As a matter of fact, many would be delighted to get a draft of equal or higher quality at machine speed and a fraction of the cost. After all, who wants to pay 10x the price and wait for weeks for a junior lawyer to manually research and type up the 20-page risk factors that are mostly based on precedents?

In these use cases, the human trust is not preserved by whether AI is used, but by whether the content generated is of good quality, and whether a human or organization would be accountable if the content is bad.

Unfortunately the watermarking technique (SynthID-Text) adopted by the leading AI companies does not convey work quality or establish human accountability.

To understand why, let’s take a look at how Anthropics explains how watermarking works. AI builds sentences one word at a time. When AI is writing the sentence “The weather today was cold and…”, the next word could be “overcast” or “grey.” Because both words work perfectly well, the AI usually rolls a digital die to pick one at random. Watermarking changes this random die-roll into a fixed sequence, a secret “key” to decide whether the AI picks “overcast” or “grey.” The same “key” is used throughout the writing, forming a hidden sequence based on the word choices. Because “overcast” and “gray” both make sense, the quality of the writing does not change. But one draft (e.g., the draft with “overcast” and other words decided by AI using the sequence) would be labeled as AI-generated by the company who holds the “key”.

This technique has several limitations:

  • No Signal between AI-Written or AI-Assisted: Anthropic acknowledged that a watermark can only determine that Claude was likely involved with the content at some point. It cannot distinguish “Claude wrote this” from “Claude heavily edited this.” As a result, a well-thought-out draft based on human-created ideas could be marked as AI-generated, while “AI slop” could evade detection if enough words are swapped with equally meaningless words.

  • No Interoperability: Detection built by Anthropic only knows how to check for Claude-generated text, and Google for Gemini-generated text, because one does not have the other company’s key. Theoretically, one could evade detection by blending different sentences generated by different proprietary and open-weight models into one single draft.

  • No Accountability Signal: As put by Anthropic, “There’s nothing in the watermark, or its key, that would allow anyone to recover any information about the user, their organization, or their chats with Claude.” Although this is good for privacy, it does not offer any indication of who drafted the text, who verified it, or who is accountable for its quality or accuracy.

  • Overlabeling Creates Noise: Anthropic acknowledged that translation would be marked as AI-generated because every word is selected by the model, so would original text that is subject to certain level of proofreading and editing. This over-labelling risks the opposite of human trust: it creates so much noise that the public will simply ignore it.

  • Disproportionate Burden on AI Systems using Open-Weight Models: This technique will require providers hosting open-weight models to build and maintain a bespoke watermarking system and detection API for each model they deploy, either for free or for a small fee. For these teams, this imposes significant ongoing infrastructure costs without solving the broader interoperability problem.

If watermarking does not enhance human trust, why are leading AI companies adopting it?

When faced with this question, Anthropic was candid:

“We’re implementing watermarking to comply with the EU AI Act… We’re applying watermarking globally at launch because we don’t yet have a durable way to scope it by region. However, we will continue to evaluate different approaches, and will share updates when we have them.”

In short: we did it because the regulator asked.

Whether the law is written too broadly or the technical implementation is too blunt, the prevailing technique invites misuse and risks stigmatizing AI use. As downstream providers and users of AI models, we have a choice: we can either allow a culture of AI stigma, or reinvent our path for human trust. In my next blog, I will do a deep dive into the EU AI Act’s transparency obligations and provide operational guidance on how to reinvent that path.

Before I wrap up this blog, I would like to address the “elephant in the room”: do I use AI to write my weekly blogs or prepare materials for the Stanford and Berkeley Law Executive Education courses.

The answer is: Yes, I do, and proudly so.

Here is what that collaborative workflow looks like:

  • I dictate raw thoughts into 3-4 AI tools and ask for first drafts.

  • I prompt AI multiple times to test my assumptions, sharpen core messages and cut out empty sentences.

  • I use AI to conduct deep research on legal and technical topics, and ask for citations to original sources.

  • I ask AI to summarize lengthy source documents before I read the provisions relevant to my blog/script.

  • I use AI to check for grammar, polish sentences and improve readability.

  • I use AI to record and edit videos, removing filler words and regenerating mispronounced words.

At the same time, I am mindful of the growing bias against AI-assisted content. That is why:

  • I do not use synthetic AI avatars or voices in my lectures.

  • During proofreading, I swap words and rewrite sentences to avoid overbroad watermarking.

I believe deeply in the power of AI, and I have no shortage of original thoughts to share with the world. The path forward is about finding a workable balance without losing sight of the original purpose of the law: preserving human trust.

—---------------

For more practical tips on AI governance and innovation, check out GenAI for the Legal Profession: Power User Edition, AI Strategy for Legal Leaders, Atticus AI Habits Workshop and my Fairly AI blogs.

No posts

Read the original on weichen221.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.