AI Model Watermarking Changes Agent Behavior
The Register, Thursday, September 17th, 2026
Lasso Security finds SynthID-Text watermarking alters tool calling and safety refusals, especially under prompt injection.
The EU AI Act requires providers to mark model output with machine-readable code for provenance, and Google DeepMind's SynthID-Text, adopted by Anthropic and OpenAI, does this by nudging the model toward one statistically likely next token over another.
Lasso Security reports that because the technique changes the process by which each token is generated, it can also change safety behavior, including whether a model refuses a harmful request and whether that refusal holds under prompt injection.
The altered behavior is not necessarily worse, but it can be. Crucially the effect crosses organizational boundaries: an agent built on a third-party harness or an API client calling an Anthropic model inherits whatever output variation follows from Anthropic's watermarking.