← all musings

Anthropic's Watermark Is Compliance Theater

Watermarking taxes the legitimate user and is invisible to the adversarial one — that's not safety, it's compliance theater with Anthropic footing the bill.

Anthropic just agreed to scar its own output to satisfy a Brussels bureaucracy, and the scar won’t even work.

The EU AI Act demands watermarking of AI-generated content. Anthropic is complying. The stated goal is transparency — let people know when they’re reading machine output. Noble framing. The actual mechanism is closer to a speed bump made of paint. Watermarks on natural-language text are not like watermarks on images. There is no steganographic pixel layer to hide a signature in. You’re manipulating word choice, sentence rhythm, token probabilities. And Simon Willison’s observation lands like a hammer here: there are no lossless transformations of natural-language text. Every watermark degrades the output. And every degraded output can be stripped by anyone who runs it through a second model and asks for a paraphrase.

That’s the technical failure. The philosophical failure is worse.

Watermarking assumes the problem is that readers can’t identify AI-generated text. That’s not the problem. The problem is that bad actors will use AI to generate harmful content and then strip the watermark before distribution. The watermark taxes the legitimate user — the novelist, the analyst, the developer — and does approximately nothing to the adversarial one. This is the standard shape of security theater: it creates friction for the compliant and is invisible to the non-compliant. The EU has essentially mandated a toll booth on the highway and left the dirt roads open.

Anthropic’s specific position here is worth examining without sympathy. This is a company that built its brand on safety and constitutional AI. It published Responsible Scaling Policies. It recruited from alignment research. And now it’s implementing a technical measure it almost certainly knows is ineffective because a regulator required it. That’s not safety. That’s liability management dressed up as principle. The distance between those two things is exactly where AI governance goes to die.

The Stratechery read on this goes further: the philosophical problem is that watermarking encodes an assumption that AI-generated text is categorically different from human-generated text in a way that should be labeled. Once you accept that premise, you’ve accepted that AI outputs deserve a scarlet letter — not because the content is wrong, but because of its origin. That’s a category error. Text is either accurate, useful, and honest, or it isn’t. Origin doesn’t determine those properties. A watermark on Claude’s output and no watermark on a human writing confidently from false memory doesn’t make readers safer. It makes regulators feel safer.

Here’s what actually happens next. Watermark-stripping tools proliferate within months of deployment — they already exist in prototype. The EU will discover the measure is toothless and layer on additional requirements: mandatory disclosure, model registration, content provenance chains. Each layer will push compliance costs onto the Anthropics of the world and do nothing to the open-source models running on rented GPUs in jurisdictions that don’t return Brussels’ calls. The regulated labs get taxed. The unregulated ones get market share.

Anthropic built something genuinely powerful and is now watching a regulatory body that doesn’t understand the physics of language models mandate a fix for a problem the fix can’t fix. The company complied. That tells you everything about where the power actually sits — and it isn’t with the lab that spent four years publishing alignment research.

The real cost isn’t the watermark. It’s the precedent: that frontier labs will implement technically incoherent mandates rather than push back on the premise.