Anthropic’s announcement that all Claude models will soon begin “watermarking” everything they generate, including text, to comply with an EU regulation, has led to speculation about how this will work. Initially, it was thought that invisible Unicode characters might be used, but Anthropic’s new approach involves a form of steganography where word choices at inference time leave detectable fingerprints.
Anthropic’s original support document claimed the watermark would be “imperceptible” and “doesn’t change the meaning, quality, or readability” of Claude’s responses. However, the author argues that this is precisely what the system will do, leading to an adulteration and corruption of the text’s semantics. The author’s error was in believing Anthropic’s claim that their system would not compromise the text’s quality.
Anthropic’s updated article, “How Claude’s Text Watermark Works,” explains the process in more accessible terms. The technique is compared to coin flipping: by analyzing word choices (tokens) at each decision point, a probabilistic determination can be made about whether the text was generated by a specific AI model. Words are categorized into “green” and “red” lists deterministically on the fly, making certain words slightly more likely to be chosen. The more words in the text, the more confident the analysis can be.
Objections to this technical premise include the fact that synonyms do not carry the exact same meaning, and the author believes that any tool used should prioritize clarity and precision in word choice. The idea that anything other than the user’s needs should influence text generation is considered offensive. Anthropic’s plan to adulterate all text longer than 200 tokens, even in private conversations, is seen as a significant drawback.
The author also criticizes the EU regulation motivating this change, calling it “red-tape nanny-state pipe-dream nonsense.” The regulation’s requirement for watermarking to be robust against typical processing solutions like screenshots, copy-pasting, and translations is deemed impractical. Tools like James Padolsey’s Declaude are mentioned as examples of how these watermarks can be circumvented.
Google’s SynthID system for watermarking images, video, audio, and text is also discussed. The author points out that Google’s own description of SynthID, which suggests that differences in word choices like “bananas” versus “airplanes” are not noticeable, is absurd. The author argues that the semantic difference is noticeable and that the system calls every word choice into question by introducing uncertainty about whether the choice was based on meaning or watermarking.
Anthropic’s new explanation is analyzed, with the author translating their claims into a more critical perspective. The assertion that the difference between watermarked and un-watermarked text is indistinguishable to readers is met with skepticism. The claim that watermarking won’t be specific to Claude is also challenged, as other providers have not announced similar global implementations. The argument that choices between words like “grey” and “overcast” don’t matter is seen as the core of the problem, making the process more perverse due to its sneakiness.
The author dismisses Anthropic’s reliance on Google’s SynthID-Text paper and its thumbs-up/thumbs-down data as insufficient evidence of quality. The paper’s conclusion that the difference in quality is “negligible” is interpreted as meaning it’s only slightly worse and that users are too indifferent to notice. The author also questions Anthropic’s claim of global implementation due to an EU regulation, given the company’s high valuation and potential IPO, suggesting either technical incapability or an overestimation of their control over their technology.
OpenAI’s approach to provenance signals is mentioned, with the author noting the ambiguity in their statement and suggesting that OpenAI could differentiate itself by making watermarking optional. The author concludes by referencing several research papers and commentary from other writers, emphasizing the “poisonous” nature of secrets in this context and the potential for unjust accusations of AI generation.




