Anthropic will deploy SynthID-Text in its future Claude models to comply with new European Union watermarking laws. This system uses a secret key to subtly alter word selection, changing a likely output like “cloudy” to “overcast” without human notice. The key allows verification of the text’s origin, but recent research indicates it also modifies how models invoke tools and follow safety guardrails. Under adversarial conditions, attackers can exploit these changes to bypass restrictions that previously prevented harmful actions. Instructions normally rejected by the system may now be executed when watermarking is active. Andrea Siposova at Lasso Security notes that altering generation processes inevitably creates tradeoffs that manifest in behaviour. Developers must now test how their large language models and agents perform specifically when watermarking is in place. The findings highlight a direct conflict between provenance tracking and security integrity.
- Watermarking keys can force models to ignore standard safety filters.
- Adversarial prompts succeed more often when text generation is altered.
- Security teams need new testing protocols for watermarked agents.




