A Tattoo You Can’t See and Can’t Check
Anthropic announced this month that Claude now embeds an invisible watermark in the text it generates. Every response, quietly tagged. Not visible to you, not visible to the reader, but detectable to anyone holding Anthropic’s key. Files get a lighter version of the same idea: a signed C2PA credential tucked into the metadata of any image or document Claude produces.
The stated reason is compliance. The EU’s AI Act Transparency Code took effect August 2, and it requires AI providers to mark generated content in a machine-readable way. Anthropic signed on, along with Google, Meta, OpenAI, and around 190 other companies. Fine. But the way Anthropic chose to implement this is worth pushing back on, and I don’t think the pushback is just Reddit noise.
The law is regional. The watermark isn’t.
The EU AI Act applies in the EU. Anthropic’s watermark applies to everyone, everywhere, all the time, because, in the company’s own words, they “don’t yet have a durable way to scope it by region.” That’s a technical limitation being used to justify a global default. Every Claude user outside the EU is now carrying compliance infrastructure for a law their country never passed, because building the regional switch was apparently harder than just flipping it on for the whole world.
It mostly catches the people who weren’t hiding anything
Here’s the part that actually bothers me. Anthropic’s own explanation of how the watermark works admits its limits plainly: light edits often survive it, but a full rewrite defeats it completely, and it’s arguable whether text that’s been rewritten word-for-word even counts as “AI-generated” anymore. So the watermark reliably tags a journalist’s transcript summary or a writer’s paraphrase request, the low-effort, honest uses. It’s already failing to tag the outputs of anyone determined enough to run their text through a rewriter first. Within days of the announcement, a GitHub project claiming to strip Claude’s watermark had over 4,500 stars, alongside a handful of other tools making the same promise. Nobody can verify those tools actually work yet, because Anthropic hasn’t shipped a detection API. But the arms race started immediately anyway, which tells you exactly who a determined bad actor expects to outrun and who they don’t.
Assistance and authorship aren’t the same thing, and the watermark can’t tell them apart
Ask Claude to reorganize a paragraph, and if enough of the resulting words were its choices rather than yours, the watermark can attach to it. Anthropic is upfront that the tool can only answer “was Claude likely involved,” not “did Claude write this” versus “did Claude clean up three sentences of this.” A tool sold on the premise of transparency ought to be more precise than that. Instead it treats “I asked for synonyms” and “I asked it to write the whole thing” as roughly the same signal, which means the anxiety it’s generating (and the Reddit threads are full of it) isn’t really unreasonable. People who used Claude for something minor are being told there’s now a mark on their work that a future tool, not yet built, might read as evidence of more than there was.
To be fair to Anthropic
Their own write-up on this is unusually thorough, and a few of their claims hold up: the watermark doesn’t cost extra tokens, doesn’t slow the model down, and by design can’t be traced back to a specific person or account. Those are real, meaningful constraints they put on themselves, and worth crediting. This isn’t a company hiding what it built or how it works.
But “we built it responsibly” and “this was the right thing to build” are two different arguments, and only the first one has been made well.
Why this matters
A watermark that’s easy to strip with enough effort, that can’t distinguish authorship from a light edit, and that applies globally because scoping it was inconvenient, isn’t really a transparency tool. It’s a policy compliance checkbox that happens to be invisible, wrapped in language about honesty that doesn’t hold up once you look at who it actually catches. If Anthropic wants this to mean something, the fix isn’t more explainer pages defending the design. It’s a version that’s regionally scoped like the law that prompted it, and a detection tool that ships before the anxiety does instead of months after.
