Anthropic has found a remarkably efficient way to trigger a philosophical crisis: put an invisible watermark on everything Claude writes.
This week, the company announced that Claude will begin embedding machine-readable marks into the content it produces. Text from supported models will carry an imperceptible watermark; generated files such as images will get digitally signed provenance metadata. The policy will eventually span Claude, its API, Claude Code, Cowork, cloud deployments, and essentially every other surface where Claude appears.
The reaction has been predictably calm and measured, by which I mean the internet immediately lost its mind.
Underneath the outrage, though, are three different questions:
Why is Anthropic doing this?
How do you watermark a sentence?
And what does the watermark actually tell you?
The answers are, roughly: Europe, statistics and less than you might hope.
The immediate catalyst is regulation.
Article 50 of the EU AI Act requires providers of generative AI systems to make AI-generated or manipulated outputs machine-readable and detectable as such. Those transparency requirements became applicable on August 2, 2026. The EU’s accompanying Code of Practice explicitly contemplates marking and detection systems for generated text, images, audio, and video.
Anthropic signed the code and, rather than maintaining one Claude for Europe and another for everyone else, is rolling the requirement out globally. Brussels has once again achieved one of its signature technological accomplishments: regulating a feature into existence worldwide.
Watermarking an image is intuitive. There are pixels to manipulate and metadata to attach. Text is harder.
Anthropic has not disclosed the precise implementation of its text watermark, but modern LLM watermarking generally exploits something fundamental about how language models work. A language model does not normally decide that the next word must be one particular word. It calculates probabilities across many possible tokens. After “The market was,” for example, volatile, weak, strong, chaotic, and mixed might all be reasonable continuations.
A watermarking system can secretly divide possible tokens into preferred and non-preferred groups, then slightly bias generation toward the preferred ones. The resulting sentence still sounds normal. But across hundreds of tokens, a statistically unusual pattern emerges.
Think of Claude flipping a slightly weighted coin every time it chooses among equally reasonable words. No individual choice gives the game away. Collect enough choices and you can conclude that the coin was probably weighted. That is the basic magic: the watermark can live in the probability distribution without visibly living in the prose.
A watermark does not prove Claude wrote something, just that Claude may have processed it. A human can write an entire document, ask Claude to fix three typos or translate it, and wind up with marked output. Conversely, Claude can write something from scratch, a user can substantially rewrite it, and the watermark may disappear.
Anthropic explicitly warns about this. Proofreading, translation, summarization and file conversion can all result in marked outputs even when the underlying material came from a human. Conversely, heavy editing, paraphrasing, translation, mixing with other text, or simply using a sufficiently short passage can make the signal undetectable.
This creates a wonderfully modern epistemological problem: we are building an authorship detector that actually detects contact. The watermark only says: CLAUDE WAS HERE. Technically true. Semantically impoverished.
Reducing human-AI co-creation to an “AI-generated” label can collapse a complicated production process into a binary signal that says surprisingly little about authorship, intent or truthfulness.
Society is currently treating “AI-generated” as a vaguely radioactive category encompassing an absurdly wide spectrum of behavior: plagiarism, automation, fraud, productivity software and creative collaboration. That category is doing too much work.
Anthropic itself seems to understand this better than some of the institutions that will eventually consume its signal. Its documentation repeatedly warns that a mark indicates processing, not complete provenance, and that the absence of a mark proves little. That caveat is crucial because probabilistic technological signals have a nasty habit of becoming deterministic bureaucratic rules.
But a professor, a publisher, a recruiter only sees ‘AI detected.’ The fact that the candidate merely asked Claude to proofread is not easily watermarked. The danger is that the watermark becomes institutionally overinterpreted.
There is another irony. Watermarking is strongest against the people least motivated to defeat it.
Significant rewriting, paraphrasing or translation may destroy the signal. File provenance metadata can disappear through format conversion, screenshots or other transformations. Academic research has likewise demonstrated a long list of attacks on generative-AI watermarking schemes, from removal to spoofing and watermark stealing.
Which produces the classic security-policy paradox: the sophisticated bad actor launders the content, the ordinary employee who asked Claude to make an email sound less passive-aggressive leaves the forensic evidence intact. Watermarking may therefore be most effective at catching people who were not particularly trying to hide anything.
We are probably approaching the expiration date of AI-generated as a useful category. Nobody asks what percentage of a Morgan Stanley presentation was “computer-generated.” The computer is assumed. AI is headed in the same direction.
As models move from destination apps into operating systems, IDEs, browsers, productivity suites and autonomous workflows, practically every important digital artifact will eventually have some machine involvement.
At that point, “Did AI touch this?” becomes considerably less interesting.
What you actually want to know is: Did a human originate the argument? Was the underlying evidence real? Was the output independently reviewed? Was someone trying to deceive the reader? Can we reconstruct how the artifact was produced? Those are much better questions.
Today, an AI watermark feels vaguely accusatory: Caught you! Claude helped! That stigma will not survive widespread adoption.
I think watermarking is directionally useful, even if the current implementation is somewhat confused.
For images, audio and video, provenance is enormously important. In a world where a convincing photograph or recording can be manufactured for pennies, having cryptographic evidence about where a file originated and how it was modified is valuable infrastructure. For text, the signal is weaker because authorship itself is becoming more fluid.
The more durable end state is not a giant binary distinction between HUMAN and AI. It is something closer to a nutrition label for information: provenance describing which systems touched an artifact, what transformations occurred, whether humans reviewed it, and how much of its history can be cryptographically verified. The future is less AI detection and more content lineage.
The watermark controversy is really a society discovering that AI has broken our clean binary between tool and author.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.