RSS Amplifier

Prof. Ryan Kim · Aug 17, 2026

Can Claude’s Watermark Prove AI Authorship?

0
Sign in to vote or save

Ryan SangBaek Kim · Prof. Ryan Kim

Suppose I write a research article from beginning to end.

The question is mine. The theory is mine. I choose the evidence, build the argument, reject alternative explanations, structure the paper, and take responsibility for every claim.

Then I ask Claude to translate it into English.

Who wrote the paper?

That question is no longer hypothetical.

On August 14, Anthropic explained how its new text watermark works. Claude will embed an imperceptible statistical signal into supported generated text. The watermark contains no hidden characters and no identifying information. Instead, it changes the source of randomness used when Claude chooses among linguistically plausible next words. Over enough text, those tiny choices accumulate into a pattern that can later be detected with the corresponding key. Claude uses a version of Google DeepMind’s SynthID-Text approach.

And Anthropic makes one particularly important point:

A translation produced by Claude carries the watermark because Claude chooses every word of the translated text.

Now place that fact beside European law.

The European Commission’s July 2026 guidance on Article 50 of the AI Act treats standard editing and certain non-substantive transformations differently from genuinely generated or materially altered content. Its treatment of translation is especially revealing. The Commission identifies AI-generated translation as a case that can fall outside the provider-side machine-readable marking requirement, and its human-review discussion uses an AI-assisted translation of a human-written article as an example of content that can qualify for the editorial-review exception. Article 50’s transparency obligations have applied since August 2, 2026.

So we now have an unusual situation.

A regulatory framework can treat a reviewed AI translation as sufficiently continuous with the human original that mandatory marking is unnecessary.

Yet the technical provenance signal may be strongest precisely because the model has selected nearly every word in the target language.

This is not a contradiction in the law. The AI Act sets legal obligations; Anthropic is free to watermark more broadly, and Anthropic says it is deploying the system globally because it does not yet have a durable way to scope watermarking by region.

But the mismatch reveals something important.

The law and the watermark are measuring different things.

And we are in danger of forgetting that.

A text watermark is easy to misunderstand because the word watermark suggests something added after the fact, like a logo stamped onto a photograph.

That is not what is happening here.

Large language models generate text sequentially. At some positions, only one continuation makes much sense. If Claude writes “Isaac Newton’s most famous work was the Principia…,” there is very little freedom in what comes next. At other positions, several words may be equally reasonable.

Watermarking exploits those low-stakes choices.

Anthropic describes a system in which the secret key and preceding context influence the source of randomness used to select among acceptable candidates. No strange word needs to be inserted. No invisible character appears. The finished sentence can remain perfectly ordinary. But across hundreds of choices, a keyed statistical pattern emerges.

This immediately explains several limitations.

Short passages carry less evidence.

Highly constrained factual language carries less evidence.

Code often carries less evidence because many tokens cannot be freely substituted without breaking the program.

Proofreading may carry almost no detectable signal because most of the words remain the human writer’s.

Translation is different.

The intellectual content may remain stable, but the target-language realization is generated nearly from scratch.

That distinction is the key to the entire problem.

Consider two variables.

The first is model control over linguistic realization: how much of the final wording the model actually selects.

The second is human intellectual contribution: who formed the research question, supplied the conceptual architecture, chose the methods, made the inferential decisions, and assumed responsibility for the claims.

Modern scholarly publishing already knows that contribution is multidimensional. The CRediT taxonomy distinguishes roles such as Conceptualization, Methodology, Formal Analysis, Writing – Original Draft, and Writing – Review & Editing. It does not reduce authorship to keystrokes.

A text watermark, by contrast, has no access to those roles.

It observes a statistical trace in linguistic production.

That suggests a hypothesis worth taking seriously:

Watermark strength and intellectual authorship need not increase together.

Translation offers an unusually clean case.

A researcher can retain virtually all of the conceptualization, methodology, analysis, structure, judgment, and responsibility while giving the model enormous control over the surface language.

High model involvement in wording can coexist with high human intellectual contribution.

Now reverse the situation.

A model may contribute substantially to an argument, outline, or conceptual structure, after which the wording is extensively reconstructed. Anthropic itself acknowledges that a complete rewrite can eliminate its watermark.

In one workflow, strong human authorship may coexist with a strong watermark.

In another, substantial AI contribution may coexist with no detectable watermark from the original model.

That does not prove the two variables are statistically orthogonal. We need data for that.

But it gives us a much better research question than “Can the detector catch AI?”

The question is:

What is the empirical relationship between provenance-signal strength and the dimensions of contribution that institutions actually care about?

I do not think we know the answer yet.

Much of the technical debate has focused on robustness.

That debate matters. Research has shown that text watermarks can face scrubbing, paraphrasing, watermark stealing, and even spoofing. Jovanović, Staab, and Vechev demonstrated that black-box access to watermarked model outputs can, for several watermarking schemes, support both removal attacks and the creation of misleading watermark-like signals. Susser, Thickstun, and Vidan have recently emphasized that these are not only engineering problems but evidentiary ones.

The correct lesson is not that “watermarks are easy to defeat.”

Some schemes are more robust than others. Claude’s actual detector has not yet been publicly released, so claims about its error rates or attack resistance would be premature. Anthropic says a detection API is forthcoming and that implementation details are still being worked out.

The more general lesson is this:

Robustness is always robustness against a specified class of transformations.

A signal can survive light editing and fail under a full rewrite.

It can be strong in unconstrained prose and weak in code.

It can be easy to detect in long text and statistically thin in short text.

The EU’s own implementation framework recognizes this technical reality. The Code of Practice treats free-form text differently from media that can carry richer provenance metadata and recognizes a roughly 200-token boundary for what counts as “very short text” under the current state of the art. It also allows access to less reliable free-form text detection to be limited to verified expert users rather than exposed indiscriminately.

That is a remarkably important regulatory admission.

The detector is not being treated as a magic truth machine.

Neither should universities, publishers, employers, or courts.

This is the point I think deserves more attention.

Imagine that Anthropic solves every technical problem.

Imagine a watermark that cannot be removed.

No false positives.

No false negatives.

No spoofing.

The detector tells us with perfect accuracy that Claude processed a particular passage.

Would that settle authorship?

No.

It would settle the fact it was designed to settle.

Claude was involved.

That is not the same proposition as:

Claude conceived the argument.

Claude determined the structure.

Claude supplied the substantive contribution.

Claude is responsible for the claims.

Claude is the author.

Anthropic itself draws this boundary. Its documentation says watermark detection cannot distinguish between Claude writing a text and Claude heavily editing one, and that the watermark says nothing about ownership or authorship.

Luciano Floridi recently made a related point in Philosophy & Technology. Publishing has become preoccupied with provenance tools, but provenance is the wrong test for authorship or quality. He argues instead for answerability: who can answer for the claims and decisions embodied in a work.

That distinction is essential.

But the new Claude system allows us to push the question one step further.

The problem may not simply be that provenance is an imperfect proxy for authorship.

The two may respond to fundamentally different features of the writing process.

A technically excellent provenance measure can therefore become a poor authorship measure not because it malfunctioned, but because we asked the instrument to measure the wrong human variable.

That is a measurement problem before it is an ethics problem.

Now the problem becomes more interesting.

Suppose a university adopts a rule:

Essays with a detectable AI watermark receive additional scrutiny.

Or a journal tells authors:

Watermark detection above a threshold requires an explanation of AI use.

Or an employer begins screening writing samples in the same way.

At that moment, the detector no longer observes an independent world.

Writers learn that it exists.

And writers adapt.

They change which tools they use.

They change where in the workflow they use them.

They alter the division between translation, drafting, editing, and rewriting.

They may deliberately retain some forms of AI assistance and avoid others.

The deployment of the measurement changes the behavior being measured.

Machine learning already has a language for this. Perdomo and colleagues call the broader phenomenon performative prediction: once a prediction or classification enters the world and affects decisions, people and systems respond, shifting the future data distribution on which the prediction operates.

Watermark-based authorship policing creates a closely related possibility.

Before institutional deployment, we might ask:

How strongly is watermark detection associated with a given AI-assisted writing workflow?

After deployment, the relevant question becomes:

How strongly is watermark detection associated with that workflow after writers know the consequences of detection and have changed their behavior accordingly?

Those are not necessarily the same distribution.

And this leads to a more unsettling possibility:

The evidentiary value of a watermark may be endogenous to the institutions that use it.

The harder an institution relies on the mark as an authorship screen, the stronger the incentive to reorganize writing practices around what the mark can and cannot see.

As those practices change, the relationship between watermark presence and substantive AI contribution can change with them.

A detector can therefore become less informative about the thing an institution cares about precisely because the institution has made the detector consequential.

This is not an argument against watermarking.

It is an argument against treating detector validity as a permanent property established once in a benchmark.

It would be easy for an American professor, editor, lawyer, or researcher to treat Article 50 as a Brussels problem.

That would be a mistake.

Anthropic explicitly says it is introducing the watermark globally at launch because it cannot yet durably scope the feature by region.

European transparency regulation is therefore influencing the statistical structure of text produced by Americans using Claude.

A doctoral student in Boston.

A lawyer in Chicago.

A novelist in Brooklyn.

A Korean scientist translating a paper for an American journal.

They can all encounter the same mark.

This makes the governance question larger than formal compliance with the AI Act.

The question becomes what downstream institutions decide the mark means.

A watermark can have one technical definition at the model provider and acquire a much stronger social meaning at the university, journal, company, or courtroom.

That gap is where mistakes become institutional.

My own research has focused on a related problem in another domain: what happens when machine interpretations acquire institutional standing over a person’s own account.

Emotion AI provides a clear example. A system may infer a person’s affective state, and that inference can become consequential when a school, employer, or other institution treats it as the operative interpretation.

Authorship is not identical to emotion.

A writer does not have infallible first-person authority simply by saying, “I wrote this.” Ghostwriting exists. Collaborative authorship exists. People can misrepresent their contribution.

So the principle cannot be that human testimony automatically defeats machine evidence.

The narrower principle is stronger:

A watermark should have evidentiary standing, not adjudicative standing.

A provenance signal should enter an evidentiary record alongside drafts, version histories, contribution statements, editorial records, and the writer’s account of the workflow.

It should not silently become the default verdict from which the human must prove innocence.

That distinction will matter increasingly in academic integrity, publishing, intellectual property, employment screening, and litigation.

We do not need another seven-point AI ethics checklist.

Three rules would already improve the situation.

First, institutions should demand evidence about the detector, not merely evidence from the detector.

If a watermark is used in an individual decision, institutions should know its dependence on text length, linguistic entropy, editing, translation, relevant transformations, calibration, and known spoofing risks.

A probability without its measurement conditions is not evidence. It is decoration.

Second, watermark detection alone should not reverse the burden of proof.

“Claude was probably involved” is not equivalent to “the author committed misconduct.”

The difference becomes especially important when legitimate workflows such as translation, language assistance, and editing can produce very different watermark intensities without corresponding differences in intellectual contribution.

Third, evidentiary validity must be tested after deployment.

If a university or publisher makes watermark detection consequential, it should periodically reassess the relationship between the signal and the behavior it intends to detect after users have adapted to the rule.

Pre-deployment accuracy is not enough for a performative environment.

Anthropic has not yet released the detection API publicly. When research access becomes available, there is a straightforward experiment worth running.

Take the same underlying intellectual work and manipulate two dimensions.

The first is intellectual contribution: human-led versus model-led conceptual and argumentative work, mapped using established contribution categories rather than an invented authorship score.

The second is surface realization: how much of the final language Claude actually generates.

One crucial condition would involve a fully human-conceived article translated by Claude.

Another would involve substantive model contribution to the argument followed by extensive human rewriting.

The texts should be length-matched and comfortably exceed the current short-text threshold.

Then compare watermark strength with the underlying contribution structure.

The important hypothesis is not that the correlation equals exactly zero.

It is more specific:

Watermark strength should respond strongly to who realizes the surface language, while responding much less reliably to who contributed the intellectual substance.

A second study could expose participants to different institutional detection policies and measure which writing workflows they say they would choose.

That would test the first step in the performative mechanism: whether the rule itself changes the distribution of behavior.

These are empirical questions.

They should be treated as such.

Watermarking may become one of the most useful provenance technologies of the generative AI era.

It may help researchers estimate synthetic-content prevalence.

It may assist platforms with large-scale transparency.

It may help investigators reconstruct parts of a production history.

But every measurement technology creates a temptation.

Once we can measure something cleanly, we begin to treat the measured variable as the variable we cared about all along.

That is where caution is needed.

A watermark may someday tell us, with extraordinary accuracy, that Claude passed through a text.

It still cannot tell us, by itself, who conceived the argument, who made the decisive intellectual choices, who rejected the alternatives, or who must answer for the claims.

Those are different questions.

And if we turn the first signal into an institutional answer to the second question, the danger will not be that the technology failed.

The danger will be that it worked exactly as designed.

A technically accurate signal can still be aimed at the wrong human question.

Anthropic. “How Claude’s text watermark works.” August 14, 2026.

European Commission. Guidelines on the implementation of the transparency obligations for certain AI systems under Article 50 of the AI Act. C(2026) 5054 final, July 20, 2026.

European Commission. Code of Practice on Transparency of AI-Generated Content. 2026.

Dathathri, S. et al. “Scalable watermarking for identifying large language model outputs.” Nature 634, 818–823 (2024). DOI: 10.1038/s41586-024-08025-4.

Floridi, L. “AI and the Future of Publishing: Not Detection, but Answerability.” Philosophy & Technology 39, 129 (2026). DOI: 10.1007/s13347-026-01148-8.

Susser, D., Thickstun, J., & Vidan, G. “From Forensics to Ecosystems: Rethinking Watermarks for Generative AI Oversight.” arXiv:2608.07337 (2026).

Jovanović, N., Staab, R., & Vechev, M. “Watermark Stealing in Large Language Models.” Proceedings of ICML 2024.

Perdomo, J. et al. “Performative Prediction.” Proceedings of ICML 2020, PMLR 119, 7599–7609.

CRediT. Contributor Roles Taxonomy, NISO.

Read the original on profryankim.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.