Anthropic's 'Watermark' Text Adulteration in Claude Is a Perversion of Writing
Anthropic announced that all Claude models, worldwide, would begin watermarking generated text to comply with an EU regulation. The initial support document, titled "How Claude Marks AI-Generated Content," provided no technical explanation and claimed the watermark is "imperceptible" and "doesn't change the meaning, quality, or readability" of responses. The author initially speculated the company would hide invisible Unicode characters, based on these assertions.
The actual technique, later explained in a separate Anthropic post, is a form of steganography: at inference time, the model's choice of words or token outputs leaves probabilistic fingerprints that can later be detected. This method inherently alters the semantic surface of the text, contradicting the earlier claims of imperceptibility and unchanged quality.
The author argues this adulterates and corrupts the semantics of AI-generated text, calling the original document misleading for not explaining the mechanism while implying zero impact. The piece is a critical rebuttal of Anthropic's framing, suggesting the watermarking approach is a perversion of writing rather than a harmless tagging system.