Anthropic Reveals How Claude's AI Text Watermarking Works, Featuring Tamper-Resistance and Code Impact Details
By admin | Aug 15, 2026 | 3 min read
Anthropic released a blog post on Friday addressing some of the key questions surrounding how it plans to watermark text generated by its AI chatbot, Claude. The post covers topics like how the watermarking system works, whether it can be removed through editing, and what impact it has on code. This comes after the company announced earlier this week that it would implement watermarking to align with the EU AI Act’s Transparency Code, which mandates that AI companies use systems capable of identifying AI-generated content. The move has sparked debate among Claude users—on Reddit, one user called it a conspiracy against innocent users, while another argued, "The only reason you wouldn’t want this is to lie to people." Meanwhile, Business Insider reports that "dozens" of users on X have claimed to cancel their Claude subscriptions in response.
Anthropic’s post begins with a broad explanation of the watermarking concept. It describes how, when Claude makes "low-stakes choices"—such as picking between words like "overcast" and "grey" to describe the weather—it can embed a pattern in its responses that is "undetectable to the reader, but is detectable to anyone who has a key that encodes it." The company emphasized that "watermarking does not impact the quality of Claude’s output," adding that "to a reader, a watermarked response is indistinguishable from an unwatermarked one."
More specifically, Anthropic said it will adopt the SynthID-Text method introduced by the Google DeepMind team in 2024, and it plans to release a watermark detection API. The company also clarified that watermarking is different from AI detection tools offered by companies like Pangram, which analyze writing for stylistic "tells" (such as the phrase "this isn’t [X], it’s [Y]") to flag AI usage. As Anthropic put it, "Picking up on these patterns is fundamentally different from checking for a watermark."
As for whether someone could simply rewrite the text to remove the watermark, Anthropic acknowledged it’s possible, but noted that "light editing probably won’t remove the watermark completely," while "a complete rewrite where every word is replaced will." The company added, "In the latter case, of course, it’s arguable whether the text can any longer be described as AI-generated."
When it comes to text that Claude has only proofread or lightly edited, Anthropic said the watermark’s presence will depend on "the length of the text and how heavily Claude has edited it." If the AI’s involvement is minimal, "nearly all the words" will have been written by the human author, leaving "very little (if anything) for the watermark to attach to."
For code, the watermark will be less prominent than in other types of text, since the model must generate functional code and lacks the freedom to choose among equally valid alternatives. However, Anthropic noted that "in areas where there is an arbitrary choice between particular words or terms within the code, the watermark can be used, such as comments within code." Still, the company stressed that "by definition, it will have a negligible effect on the actual code produced."
Finally, Anthropic pointed out that Claude won’t be the only AI chatbot generating watermarked text. As the company explained, "other major model developers have signed the same Code of Practice and will be implementing their own watermarks."
Comments
Please log in to leave a comment.
No comments yet. Be the first to comment!