Understanding the Impact of LLM Watermarking on AI Agent Behavior
Anthropic recently announced that future Claude models would embed an invisible watermark in their output, and it later disclosed that the watermark is based on Google DeepMind’s SynthID-Text. Text watermarking itself is not new, but its deployment now has regulatory relevance: Article 50(2) of the EU AI Act requires providers of AI systems that generate synthetic text to mark outputs in a machine-readable format and make them detectable as artificially generated or manipulated, using technical solutions that are effective, interoperable, robust, and reliable as far as technically feasible.
Although watermarking is designed for provenance, SynthID-Text changes the process by which the model generates each next token. At the model level, this can change safety behavior, including whether the model refuses a harmful request and whether that refusal holds under prompt injection; at the agent level, the same sampled tokens can determine which tool is called and what arguments are passed to it.
Prompt injection connects these two settings because a weakened refusal becomes more consequential when the model can also act through tools. Such a watermarking procedure can therefore affect both what the model says and what an agent does—an effect the authors call sampling drift.
Whether this drift appears in practice is an empirical question. The authors find that it does, in both model refusal behavior and agent tool calling. The effect is model- and key-dependent and can be obscured by aggregate scores when changes in opposite directions cancel.
They therefore report both net performance and paired disagreement between watermarked and unwatermarked runs. The closing section discusses what this means for AI safety and security and what developers should do about it.
Under the heading “Built for Content Provenance, Deployed Inside Agents,” the article explains that a text watermark embeds a signal that allows output to be identified as AI-generated. Existing approaches include post-processing methods and methods integrated directly into LLM generation; generation-time approaches include logits-biasing methods, distortion-free keyed sampling, cryptographically motivated constructions, and SynthID-Text’s Tournament sampling.
Figure 1 contrasts this process with ordinary sampling: ordinary generation samples from the model’s token distribution, while a generative watermark adds a random seed generator, sampling algorithm, and scoring function, with SynthID-Text using tournament sampling (adapted from Dathathri et al.). The authors use SynthID’s non-distortionary configuration, which preserves the original token distribution in expectation over the watermark randomness while individual generations under a fixed key can still differ. Dathathri et al. report no measurable quality degradation across nearly twenty million Gemini responses.
Anthropic’s deployment also illustrates why this matters beyond first-party chat applications, according to the excerpt, though the excerpt ends at that point.