How Claude Marks AI-Generated Content
Anthropic will embed watermarks in the text and files generated by future models it launches in the EU, as part of its effort to comply with content and transparency rules in the bloc’s AI Act.
[…]
The move may further amplify the appeal of open weight models and alienate Claude customers, who don’t necessarily want consumers of their AI-generated content to know its provenance.
[…]
Anthropic itself is already hedging about the utility of its marking method, noting that detected marks are not conclusive evidence that Claude produced the content and that the absence of marks cannot guarantee that AI wasn’t involved in the creation of a particular piece of content.
Claude models launched in the EU on or after August 2, 2026 will support machine-readable marking at launch. Generated text will carry embedded watermarks, and generated files will include digitally signed provenance metadata where supported.
[…]
Marks will apply to output from supported Claude models across Claude Platform (API), Claude, Claude Code, Claude Cowork, and Claude Tag, and wherever Claude is offered, worldwide. Some platforms or features may not support certain marking types.
[…]
We’ll support users and other third parties to detect Claude’s marks, as the Code requires, and we’ll share details in forthcoming documentation.
Anthropic is claiming this is for compliance with an EU law but that “Marking will apply to output from supported models wherever Claude is offered, worldwide.”
This is infuriatingly opaque. I won’t speak to the “signed provenance metadata” they say they’ll be embedding in generated files like PDFs, SVGs, PNGs, and JPEGs. That’s important but it’s complex. Plain text is simple. A string of text is just one character followed by another. I presume that Claude is going to start “embedding” invisible characters/byte sequences between visible characters? What they’re claiming to do here seems impossible, frankly, and anything they attempt to embed invisibly is going to cause immediate problems.
[…]
If the “watermark” is not comprised of invisible characters but rather visible ones, how in the world does this jibe with their claim that it’s “imperceptible”? (This Reddit thread claims that’s how it will work — “It’s a form of steganography where the model subtly biases its word choice to create a statistical pattern that can be detected later.” How in the world can that be squared with “it doesn’t change the meaning, quality, or readability”?) And any sort of semantic detection like this is going to cause false-positive problems.
Previously:
- Hiding Vulnerabilities in Source Code
- Hiding Data in Emoji
- A Vatican-Sized Flag Mystery
- Unicode and Copying and Pasting Code
- Hiding Easter Eggs in Maps
- Genius Accuses Google of Copying Its Lyrics Data
- Fingerprinting Swift Code Using Spacecrypt
- Xerox Scanners and Photocopiers Randomly Alter Numbers
Update (2026-08-17): Anthropic provides more details:
Watermarking won’t be specific to Claude. As of August 2, the EU requires AI providers serving its market to mark AI-generated content. Other major model developers have signed the same Code of Practice and will be implementing their own watermarks.
[…]
In internal testing, we’ve seen no impact of watermarking on the content, level of creativity, or readability of Claude’s text. In the SynthID-Text paper, which introduced the technique we use, Google DeepMind tested this impact by serving a model that used watermarking to a portion of their Gemini traffic and comparing thumbs-up and thumbs-down ratings. They found no statistically significant differences from the unwatermarked model. And in a controlled study, human raters comparing watermarked and unwatermarked answers side-by-side saw no difference in quality.
[…]
There are limitations to the effectiveness of watermarking. Using our key, one can only answer the question “What is the likelihood this was partly written by Claude?” It doesn’t confirm whether the text was human-written, and it can’t tell whether the text was written by a different AI (even if that other AI uses watermarking, it would have a different key; it might also use a different watermarking method altogether). Detecting a watermark also doesn’t work well on small samples, where there are fewer word choices and thus less information to go on. As a passage increases in length, confidence about Claude’s involvement increases too.
Watermarking is sparser on factual passages where there are fewer choices that can be made without decreasing the accuracy of the text. […] For the same reason, code—which in very many cases has to be exact—has generally less watermarking than some other forms of text.
John Gruber (Mastodon, Bluesky, Hacker News):
I initially guessed “invisible characters” not because I didn’t think of the semantic word-choice technique, but because I was a fool who took Anthropic at its word in their description of what they would do.
[…]
They say “imperceptible” and “doesn’t change the meaning, quality, or readability”. Their words. Not almost imperceptible. Not slightly changes the meaning, quality, or readability.
There seems to be a slight of hand here, where Google studied whether humans saw a difference in quality but then Anthropic claimed that there was no difference in meaning. Changing words will almost always change the meaning. Perhaps not in an important way, but a change is a change. Pretending otherwise is offensive to a writer. The response to this line of argument seems to be that LLMs already have randomness and that the marking changes will be within the level noise that’s already there. To me, this presents a pessimistic view of the capabilities of LLMs, contrary to Anthropic’s normal claims. The Google study was published in 2024 and didn’t test Claude. The models have changed a lot since then, but they are assuming that the precision hasn’t advanced enough to be affected, nor will it in the future. Or that when there is a tension between watermarking and quality they will pick the former. And perhaps you won’t be able to tell because everyone else will be watermarking, too.
But the very best description of the general idea behind the technique is an interactive essay by James Padolsey, “How AI Text Watermarking Works”.
[…]
This isn’t just about text one might generate with the intention of passing it off as their own natural work. This isn’t even about LLM proofreading of work written by hand. Anthropic is saying that all new Claude models are going to adulterate every single bit of text longer than 200 tokens (~150 words) they generate, including everything it presents to its users to read. So even in a private conversation between a user and Claude, which will never be read by anyone other than the user, Claude will begin making word choices in the name of marking its output in statistically predictable ways rather than maximizing clarity and precision.
[…]
Taken literally, compliant LLM terms of service must forbid users from rephrasing the output from models that comply with this regulation, because the word choices are the marks. But it’s not the European Union that is trying to impose their absurd, impractical, witch-hunt-fueling regulation on the entire world. That falls on Anthropic.
Complying with this, particularly with regard to text, is only going to create problems for honest users. Dishonest users attempting to pass off AI-generated text as their own writing (students, employees, whoever) will simply circumvent detection through non-compliant AI paraphrasing tools.
I don’t love that checking for a watermark requires submitting the text to the AI provider, and all you get back is an opaque result with no way to verify it.
James Padolsey (via John Gruber):
The move is Anthropic’s response to Article 50(2) of the EU AI Act, which requires providers to ensure that outputs are “marked in a machine-readable format and detectable as artificially generated or manipulated”.
[…]
The Act exempts standard editing, but many legitimate assistive uses require more substantial rewriting while leaving the ideas, judgment and responsibility with the human. The same thought that led to this law could have applied to calculators at the time of their inception, had their outputs revealed themselves through artefacts. Thankfully, a sum borne of the brain is treated no differently from one produced by a calculator. Likewise with spellcheckers. To make assistance suspect only once the tool becomes capable enough to compose a whole sentence is not a principled boundary. It is a moral premium placed on difficulty itself.
Anthropic has nevertheless chosen a blanket, model-level implementation that appears broader than the law’s minimum requirement. That may be convenient compliance engineering, but it discards distinctions the law expressly attempted to preserve. The result is a signal broad enough to implicate harmless and assistive use, yet fragile enough to be removed by a motivated person through substantial recomposition. It risks concentrating suspicion on ordinary and assistive users while remaining weakest against deliberate deception.
Update (2026-08-20): See also:
- Supplementary Information for the SynthID-Text paper
- Jason Anthony Guy
- Daniel Jalkut (Mastodon)
- John Gruber (Mastodon)
- Greg Titus
- Allie
Without getting into the weeds, I still think Anthropic’s claim of their being no change in the meaning, quality, or readability is not accurate, using the plain meanings of those terms. If you read the papers, there are ways of ensuring that the frequencies of individual tokens or sequences of them are not affected in aggregate. But this is a different claim, and Anthropic does not specify which tradeoffs they’ve chosen there. I’m willing to grant that they can make the effect very small.
At this point, I think the most interesting question is, what will be the effects of these watermarks?
AI watermarking is security theater. There is absolutely nothing to prevent people from removing any possible watermark in text, images, video or other data. There is no way to prevent false positives, or false negatives, in detectors, either. It’s all superficial nonsense.
I find it interesting to be able to test a particular passage of text, but what am I supposed to do with that information? Even without watermark removal, a negative for a given model doesn’t tell you whether another model was used. And I’m uncomfortable with convicting someone based on a watermark given that false positives are possible and independent verification is not.
I asked Gemini about the benefits, and it said:
Regulatory Compliance and Policy
- Legal adherence: Helps major tech platforms satisfy emerging mandates requiring automated systems to label machine-generated media.
- Ecosystem tracking: Gives platforms and enterprises a high-level telemetry signal to monitor the spread of synthetic content across public channels.
Triage and Risk Mitigation
- Mass moderation: Assists automated filters in flagging bulk, unvetted bot content or spam networks where statistical trends matter more than single-file perfection.
- Provenance auditing: Provides an extra layer of metadata for digital media archives, helping investigators spot obvious clusters of AI-generated assets.
In other words, it’s not so much about individual pieces of content.
Update (2026-08-24): John Gruber:
Also, more and more, I do think this has little to do with the EU regulation and everything to do with Anthropic and Google wanting to watermark the text that Claude and Gemini generate for their own purposes, to identify text to avoid training future models on to avoid model collapse/AI inbreeding.
See also: Accidental Tech Podcast.
30 Comments RSS · Twitter · Mastodon
I don't understand how Gruber can have such strong opinions on things he clearly doesn't understand at all.
All Gruber-bashing aside (deserved, if only because @Michael_Tsai included it in his post), a question for all:
Since when is an *invisible* watermark considered one? If a tree fall in a forest and nobody is around to hear it, does it make a noise? Like it or not, a group of human beings (the EU) made a decision and Anthropic decides their answer is something only a non-human machine can see. Huh?
I don’t like this generic Gruber bashing. If you think he’s wrong, how about explaining where you disagree?
When I first saw this story I had the same thought as Gruber - how can you possibly add a "watermark" to text? If the watermark is detectable, it can be easily removed. This would be just like adding DRM to plain text, seems about as likely as a perpetual motion machine as far as I can see.
Just ask Claude. It links to https://www.techtimes.com/articles/323873/20260811/claude-now-watermarks-text-everywhere-mark-proves-processing-not-authorship.htm and describes a technique published by John Kirchenbauer et al.
> Roughly: the sampler is biased toward a pseudorandomly-chosen subset of the vocabulary at each step, seeded from preceding tokens, so a detector holding the key can measure a statistically improbable enrichment of those tokens across a passage.
So the actual tokens selected are biased. Rephrasing would break the watermark. It works better on longer texts. It seems like it would be pretty effective on something the length of a scientific paper or dissertation. Even if the author paraphrases/writes their own text for part of the document, if they didn’t rewrite the whole document, it could be statistically detectable as Claude-generated.
Would be less effective on short texts.
"I don’t like this generic Gruber bashing. If you think he’s wrong, how about explaining where you disagree?"
I disagree with this characterization. Gruber's own post is literally, "I don't know how this works, and the only ideas I can come up with are dumb, so this is dumb." It's not "generic Gruber bashing" to point this out.
How did it not occur to him to spend five minutes asking an LLM how it actually works? This is just shitty behavior, and I don't feel it's my job to be his research assistant and explain to him how something works if he can't be bothered to spend five minutes himself.
It's *his* job to do the most basic amount of research on the stuff he publishes, not mine.
I'm fairly certain they're over thinking it. A decent saying ALL THE TEXT BELOW THE LINE IS GENERATED BY MODELL NN
Followed by a line and the generated text should be enough. Then it's on people using the text to follow the law in a similar way.
Fingers could have metadata, but any publication or website using sound video, or images would have to clearly label them.
"A decent saying ALL THE TEXT BELOW THE LINE IS GENERATED BY MODELL NN"
That's not a watermark, though. Anthropic hasn't disclosed exactly how they embed a watermark, but based on Anthropic's post, Nate's comment above is almost certainly correct.
Sorry, the point I wanted to make was that I don't think EU demand watermarks. Just machine readable labels, så plain text should work?
(Guy guessing like Gruber)
But it was months since I went through the AI act (using ai).
There are rules targeted at users of AI services, and rules targeted at providers of AI services. Users are required to label their content in some cases, but providers are required to watermark their content in some cases.
Here's Article 50, Transparency Obligations for Providers and Deployers of Certain AI Systems:
https://artificialintelligenceact.eu/article/50/
Here's the part that requires watermarking:
"Providers of AI systems, including general-purpose AI systems, generating synthetic audio, image, video or text content, shall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated. Providers shall ensure their technical solutions are effective, interoperable, robust and reliable as far as this is technically feasible, taking into account the specificities and limitations of various types of content, the costs of implementation and the generally acknowledged state of the art, as may be reflected in relevant technical standards."
The only way to meet these requirements (machine-readable, content itself detectable as artificially generated, format must be robust) is by watermarking content.
By the way, the connection to the EU is, IMO, the reason Gruber didn't bother to actually figure out what Anthropic was doing and just decided that it was stupid shit.
@Plume Gruber did describe in the parenthetical how it actually works. I think it’s reasonable to wonder whether there’s a non-zero effect on meaning, and we don’t know what the false positive rate is yet.
Sure, let's look at what he said:
"This Reddit thread claims that’s how it will work — “It’s a form of steganography where the model subtly biases its word choice to create a statistical pattern that can be detected later.” How in the world can that be squared with “it doesn’t change the meaning, quality, or readability”?"
This is covered in the research on this topic. In most use cases, LLMs are already not deterministic; they don't produce consistent meaning, quality, or readability. Watermarking hooks into this existing randomness, so it doesn't necessarily produce worse output, just different output.
"And any sort of semantic detection like this is going to cause false-positive problems."
Everything has false positives. The exact false-positive rate is set by Anthropic; presumably, they'll set it low enough that it won't be an issue. I'll decide if I should get upset about it when I actually see what they do.
"If I ask a tool to generate text or suggest text for me, I expect that tool to generate the best possible word choices it can."
LLMs already don't do that. Ask ChatGPT the same thing twice, and you'll get different answers. Obviously, both of these can't be "the best possible word choices".
"Not corrupt its output for the sake of watermarking."
That's not what it does; see above.
"What happens if you copy my text and quote it in something you write, by hand? Now you’ve got a watermark in your prose, invisible characters or something, that implicates that your prose is AI-generated? Even though you didn’t use AI and just copy and pasted text from me? Madness."
Yes, if you copy a large portion of text that is AI-generated and put it into your own text, then it will come out as AI-generated. Which it should, because a large portion of it *is AI-generated.*
"They obviously need to explain exactly what they’re embedding in text if anyone is actually going to detect it, but if they explain it, anyone can simply remove it."
Yes, of course you can do that. You don't even need to know what Anthropic does; you can just read Claude's output and paraphrase it in your own words, and there will be no watermark. I don't think that's an argument against having a watermark; it just means it's not 100% reliable at detecting LLM involvement at any stage of the process. Which nobody ever thought it would be.
"And what happens, for example, if someone just OCRs Claude-generated text rather than selects and copies it?"
Nothing; the watermark survives that. Again, this is an example of Gruber just making something up and declaring it a problem without any effort to actually validate if the thing he made up in his head is real.
"This is all so stupid."
Again, Gruber plainly calling something stupid when he admittedly doesn't know what it is or what it does.
@Plume I think stupid is going too far, but assuming that the models will always be such that the marking biases are below the bounds of the inherent nondeterminism—that there will never be a tradeoff in quality—seems unwarranted.
Anthropic claims it doesn't impact quality:
"When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response."
This is consistent with research on similar systems:
"The Gemini user interface allows users to provide feedback on model responses via a thumbs-up (good response) and a thumbs-down (bad response). We analysed approximately 20 million watermarked and unwatermarked responses and computed the thumbs-up and thumbs-down rates (both as a fraction of the total number of thumbs-up and thumbs-down feedback received). We found that the thumbs-up rate for the two models differed by 0.01% (with the watermarked model being higher); and the thumbs-down rate differed by 0.02% (with the watermarked model being lower). We found both of these differences to be statistically insignificant, and well within the 95% confidence intervals."
And:
"As previously mentioned, a watermarking scheme can be non-distortionary, a property relating to quality preservation; however, the phrase and its variants have been used in the literature to mean several distinct definitions, causing some confusion. In this work, we resolve the confusion by providing clear definitions of non-distortion, from weakest to strongest. The weakest version is single-token non-distortion, which says that, on average over the random seed rt, the distribution of the output token xt generated by the watermarking sampling algorithm is equal to the original LLM distribution pLM(⋅∣x<t) (Fig. 1). Stronger versions of non-distortion expand this definition to one or more sequences of text, ensuring that on average the probability of the watermarking scheme generating a particular text or sequence of texts is the same as for the original LLM. Full definitions are provided in Supplementary Information section G.
(...) For our experiments, we configure SynthID-Text to be single-sequence non-distortionary; this preserves text quality and provides good detectability, while having some reduction to inter-response diversity. We call this configuration ‘non-distortionary SynthID-Text’ (and where not otherwise specified, ‘SynthID-Text’ also refers to this)."
@Plume This quote lowers my confidence in their assertion. If the user is only given one response, on a topic where they probably don’t know the answer (or they wouldn’t be asking), and they can choose to thumbs-up or thumbs-down, they are just indicating whether the response seems broadly acceptable. If you wanted to prove that there was no quality difference, I would want someone who actually knows the answer to see both responses side-by-side and rate which one is better.
@Plume Thanks for the explainer,I agree that there needs to be more than a preceding text to meet the requirements.
WRT the impact on text quality I don't think there's an issue, It's no where near good as it is.
"If the user is only given one response, on a topic where they probably don’t know the answer (or they wouldn’t be asking)"
Right, "probably", but some people will ask about things they know. So I agree that the signal of a single thumbs up/thumbs down is extremely weak, but they did it with 20 million messages. If there was a difference, we'd see it. I'd wager that most of the things we absolutely trust in everyday life have much weaker evidence behind them.
Note that they also offer the statistical analysis, so the lack of a difference in the empirical tests was the expected outcome.
Just found this, which I think is a pretty good explanation of how and why these types of systems work:
https://blog.gaborkoos.com/posts/2026-08-12-Text-Watermarking-for-Non-Academics/
Yes, I think this is sensible and am broadly supporting of this, so long as the vendors don't claim certainty. This can only add to the information available within what will presumably be tolerable confidence, which is all you could really want. As long as this isn't a prohibitionist slippery slope against open-weight models, it's fine by me.
"There seems to be a slight of hand here, where Google studied whether humans saw a difference in quality but then Anthropic claimed that there was no difference in meaning. Changing words will almost always change the meaning. Perhaps not in an important way, but a change is a change"
There was never a word there to change. There was always a list of words with different probabilities, and if you had prompted the LLM ten seconds later, you would have gotten a different word out of that list.
"The models have changed a lot since then, but they are assuming that the precision hasn’t advanced enough to be affected, nor will it in the future."
What do you mean by "precision"? LLMs are getting better at selecting good candidates for next tokens, but they're not getting rid of the randomness.
"And perhaps you won’t be able to tell because everyone else will be watermarking, too."
No, you won't be able to tell because the output you get is one you could also have gotten without watermarking. It's one of the possible outputs the LLM could have generated; watermarking just picked one of the possibilities with a specific statistical attribute.
I'm not sure if this whole discussion isn't just based on a fundamental misunderstanding of how LLMs work.
"The result is a signal broad enough to implicate harmless and assistive use"
There is no "implication." If you use LLMs to write text, just divulge that you did so, and if you want to, explain why. If you feel bad about using LLMs and don't want people to know that you use them, don't use them. But do not argue that the correct solution is to mislead people.
And while we're talking about "sleight of hand", what about this from Gruber:
"In a group chat, a friend of mine quoted the above, and I responded that if a chatbot wrote “My favorite tropical fruits are mango and airplanes”, I’m pretty sure I’d fucking notice."
Except later in the same article, Gruber admits this is a straw man, and that watermarking wouldn't make the LLM say that. The only example in the whole article where watermarking would generate an artifact noticeable to a human is not real, yet the whole post's point depends on the claim that watermarking would generate perceptible artifacts.
There is a 0% chance that Gruber will be able to tell the difference between an LLM with watermarking on and one with watermarking off. However, there is a 100% chance that people will attribute every perceived change in quality in Anthropic's models to watermarking. They're already blaming watermarking for the perceived lack of quality in models like Opus 5, which came out before Anthropic started watermarking.
@Plume
There was never a word there to change. There was always a list of words with different probabilities, and if you had prompted the LLM ten seconds later, you would have gotten a different word out of that list.
Yes, but unless the probabilities were all equal, introducing the bias will make it statistically pick words that would have been less likely apriori.
I’m certainly not arguing for misleading people, and I don’t use LLMs to write text, anyway.
> If the watermark is detectable, it can be easily removed.
No, I don’t think that holds true.
Call it a “digital fingerprint” if you prefer (tomato, tomahto). For example, Shazam works (surprisingly well) by trying to find patterns in music. Changing a song such that Shazam no longer detects it as the same song yet is recognizable to humans as “the same” is rather tricky.
As for Gruber, I think his judgment on these topics is clouded by a weird bone to pick on the EU, or on governments that attempt (with varying success / usefulness) to regulate tech. In this case, I’m not sure how “it’s useful to detect LLM-generated text” is even that controversial. Just ask teachers how they feel about it.
So I'm the only one who gives this whole argument redundant since the quality off LLM texts is so low that there's plenty of wiggle room without any loss of quality.
"Yes, but unless the probabilities were all equal, introducing the bias will make it statistically pick words that would have been less likely apriori."
That's true, but it's true in a way that makes it irrelevant to the quality claim. If the LLM has ten words available to choose from, and randomly (it's not actually random, but also not correlated with quality, so the effect in this case is random) removes five of them, each of the remaining five has twice the probability of being picked than before. But the average quality of these words remains the same.
"I’m certainly not arguing for misleading people"
Yeah, sorry, that was a generic "you", not a "you you".
Also, Gruber has another follow-up post where he makes contradictory claims, calls people idiots, and unironically claims that "English is the most expressive language in the world."
"the quality off LLM texts is so low that there's plenty of wiggle room without any loss of quality."
I think that is largely true. The quality of English prose written by LLMs tends to be quite low in ways that are mostly unrelated to word choice.
No, we won't. There's a 100% chance that people will whine about how much worse Claude is, because that's every second post on r/Anthropic. Nobody is going to test this objectively, and in six months, we'll all have forgotten about it.