AI watermarking marks origin not truth
When Anthropic announced it would watermark text generated by its models, the company framed the move as a step toward transparency. The EU AI Act’s transparency code, which took effect in August 2025, requires AI companies to mark synthetic content in ways other systems can identify On the surface, this sounds like a straightforward technical solution: embed a signal in the output, let detectors read it, and everyone knows what is machine-made. But the deeper story is about something else entirely. It is about the gap between what watermarking promises and what it actually delivers. That gap is where AI deception lives.
The watermark itself is a clever piece of engineering. Anthropic says it will be applied at the model level, meaning the signal travels with text no matter which product generates it. Copy and paste it, and the watermark persists. Edit it slightly, and it may still hold. The company even claims the mark will survive some editing, though it has not said how much. This is the kind of technical detail that sounds reassuring. The problem is that reassurance is not the same as truth. A watermark tells you a piece of text was generated by a specific model. It tells you nothing about whether that text is accurate, honest, or safe.
Consider what the watermark does not do. It does not tell you whether the model hallucinated a source. It does not tell you whether the model fabricated a quote or invented a statistic. It does not tell you whether the model was manipulated by a prompt injection attack, where hidden instructions in a document steer the output toward malicious ends. A watermark is a provenance stamp, not a truth detector. The EU regulation asks for transparency, and Anthropic has delivered exactly that. But transparency about origin is not transparency about content. The two are easily confused, and the confusion is precisely where deception thrives.
The history of content authentication is littered with similar gaps. The C2PA standard, which Anthropic will use for files, was designed to create a chain of custody for digital media It works well for photographs and videos, where metadata can be embedded at capture time. Text is different. Text has no capture moment, no camera sensor, no shutter click. Text is just characters, and characters can be rearranged indefinitely. The C2PA approach assumes a stable artifact. Text is the least stable artifact we have. This is not a technical limitation that better engineering will solve. It is a structural mismatch between the tool and the material.
The deeper issue is that watermarking addresses the wrong fear. The public worry is not that AI generates text. It is that AI generates text that looks indistinguishable from human thought. A watermark does not change the appearance of the text. It adds an invisible layer that only machines can read. The person reading a news article, a legal brief, or a medical note still cannot tell whether a human or a machine wrote it. They have to trust that the watermark is there, that it works, and that someone is checking it. That is a lot of trust to place in a system designed to be invisible.
Anthropic’s support page says the watermark will be present no matter which Claude product the text comes from. That includes Claude Code, the company’s coding assistant, and Claude Cowork, its collaboration tool. The scope is impressive. But scope is not the same as coverage. A watermark only matters if someone actually checks for it. There is no requirement in the EU regulation that platforms verify watermarks before publishing content. There is no mandate that social media companies strip or flag watermarked text. The infrastructure for checking is left to the market, and the market has shown little interest in policing synthetic content. The regulation creates an obligation for producers. It creates no corresponding obligation for distributors.

This is where the deception becomes structural. Anthropic is not deceiving anyone by adding a watermark. The deception is in the framing. The company presents watermarking as a solution to the problem of AI-generated content, and the EU presents it as a transparency measure. But the actual problem is not that we cannot identify AI text. The actual problem is that we cannot trust what AI text says. A watermark answers the question “Where did this come from?” It does not answer the question “Can I believe this?” The second question is the one that matters, and it remains unanswered.
The industry’s rush to watermarking has a familiar pattern. Last week, the music platform Suno said it would mark tracks created on its system. Last month, Substack partnered with Pangram to flag AI-generated content, and its CEO coined the term “Claudefishing” to describe people using AI to generate content Each announcement is framed as a victory for authenticity. Each one solves a narrow technical problem while leaving the broader trust problem untouched. The pattern is consistent: identify a symptom, build a tool, declare progress. The underlying disease, which is that AI can produce convincing falsehoods at scale, gets no treatment.
There is also the question of who gets to decide what counts as AI-generated. Anthropic says the watermark will be applied to all models released after August 2. Older models will get support later. But what about text that is edited heavily enough to break the watermark? What about text that is translated into another language? What about text that is paraphrased by a human who reads an AI output and rewrites it in their own words? At some point, the watermark disappears, and the text becomes effectively unmarked. The company has not said where that point is. It has not said how much editing is required. The uncertainty is not an oversight. It is a feature of the technology. Watermarks are designed to be fragile enough to be useful, but that fragility is also their weakness.
The most troubling part is what the watermark reveals about our relationship with AI. We are building systems that can generate text indistinguishable from human writing. We are then building systems to detect that text. And we are building regulations to require the detection. Each layer of the stack is a response to the layer below it. But no layer addresses the fundamental question of whether we should be generating this text in the first place. The watermark is a fig leaf over a much larger problem. It lets us pretend we have solved the issue of AI deception while the deception continues unchecked.
The EU AI Act’s transparency code is well-intentioned. It tries to create a baseline of honesty in a landscape that is increasingly dishonest. But the code, like the watermark it mandates, operates at the level of form rather than substance. It tells us what a text is, not what it means. It tells us where a text came from, not whether it is true. The regulation treats information as a product with a label. Information is not a product. It is a relationship between a claim and a reader. The watermark cannot mediate that relationship. It can only point to its own existence.
Anthropic has said it will clarify how much editing is needed to remove the watermark. That clarification will be useful, but it will not change the fundamental dynamic. The watermark is a marker of origin, not a guarantee of integrity. It is a tool for attribution, not a tool for verification. The gap between those two things is where AI deception lives. It is where a model can generate a plausible but false legal citation, a fabricated news story, or a convincing but baseless medical claim. The watermark will be there, marking the text as synthetic. The reader will still not know whether the content is correct.
The responsibility question remains. When a watermark fails to prevent deception, who is accountable? The model maker, who built the system that generates the text? The platform, which distributes the text without checking the watermark? The regulator, which mandated the watermark without requiring verification? Or the reader, who is left to navigate a landscape where every piece of text comes with a label that says “trust me, but not too much”? The answer is unclear, and the lack of clarity is itself a form of deception. We are being told that the problem is solved. The problem is not solved. It has only been labeled.

Sources
1. Anthropic
3. C2PA
4. Suno
5. Substack
6. Pangram
