How AI text watermarking works
Watermarking plain text seems impossible because text has no pixels or metadata that survive copy-paste, yet the marks are real and invisible. Google has watermarked text from the Gemini app and web experience since 2024 (with the API as a documented exception), and as of August 2026, new Claude models mark text at the model level, with earlier models to follow.
The mechanism lives in the choices between words. A model selects each word by rolling weighted dice over a shortlist of acceptable options. A secret key colors the shortlist into green and red words (the coloring depends on the preceding words, so the same word can be green after one prefix and red after another), then tilts the probabilities slightly toward green. The nudge is mild—red words can still win—so reading remains natural. Only the cumulative lean toward green accumulates, and only the key-holder knows which words were green where, enabling detection. This follows Kirchenbauer et al. 2023; Google's SynthID (the production system) achieves the same end with a subtler tournament-style approach.
Because the watermark lives in the choices rather than the characters, it survives copying and does not alter the text's meaning. This allows providers to verify whether text originated from a specific model, which has important implications for detecting AI-generated content and maintaining content authenticity.