How Claude Marks AI-Generated Content
Anthropic will embed watermarks in the text and files generated by future models it launches in the EU, as part of its effort to comply with content and transparency rules in the bloc’s AI Act.
[…]
The move may further amplify the appeal of open weight models and alienate Claude customers, who don’t necessarily want consumers of their AI-generated content to know its provenance.
[…]
Anthropic itself is already hedging about the utility of its marking method, noting that detected marks are not conclusive evidence that Claude produced the content and that the absence of marks cannot guarantee that AI wasn’t involved in the creation of a particular piece of content.
Claude models launched in the EU on or after August 2, 2026 will support machine-readable marking at launch. Generated text will carry embedded watermarks, and generated files will include digitally signed provenance metadata where supported.
[…]
Marks will apply to output from supported Claude models across Claude Platform (API), Claude, Claude Code, Claude Cowork, and Claude Tag, and wherever Claude is offered, worldwide. Some platforms or features may not support certain marking types.
[…]
We’ll support users and other third parties to detect Claude’s marks, as the Code requires, and we’ll share details in forthcoming documentation.
Anthropic is claiming this is for compliance with an EU law but that “Marking will apply to output from supported models wherever Claude is offered, worldwide.”
This is infuriatingly opaque. I won’t speak to the “signed provenance metadata” they say they’ll be embedding in generated files like PDFs, SVGs, PNGs, and JPEGs. That’s important but it’s complex. Plain text is simple. A string of text is just one character followed by another. I presume that Claude is going to start “embedding” invisible characters/byte sequences between visible characters? What they’re claiming to do here seems impossible, frankly, and anything they attempt to embed invisibly is going to cause immediate problems.
[…]
If the “watermark” is not comprised of invisible characters but rather visible ones, how in the world does this jibe with their claim that it’s “imperceptible”? (This Reddit thread claims that’s how it will work — “It’s a form of steganography where the model subtly biases its word choice to create a statistical pattern that can be detected later.” How in the world can that be squared with “it doesn’t change the meaning, quality, or readability”?) And any sort of semantic detection like this is going to cause false-positive problems.
Previously:
- Hiding Vulnerabilities in Source Code
- Hiding Data in Emoji
- A Vatican-Sized Flag Mystery
- Unicode and Copying and Pasting Code
- Hiding Easter Eggs in Maps
- Genius Accuses Google of Copying Its Lyrics Data
- Fingerprinting Swift Code Using Spacecrypt
- Xerox Scanners and Photocopiers Randomly Alter Numbers
3 Comments RSS · Twitter · Mastodon
I don't understand how Gruber can have such strong opinions on things he clearly doesn't understand at all.