🛡 SynthID-Text Watermarks Disrupt LLM Agents

According to Lasso Security (The Provenance Tax), SynthID-Text from Google DeepMind, being implemented in Claude, reduces tool-calling accuracy in 6 out of 7 models — a sampling drift effect. In paired runs with identical seeds, an average of 6.5% of verdicts diverged, up to 16.8% for phi-4.

🌍 Marking under Article 50(2) of the EU AI Act is no longer free: a watermark at the sampling stage can replace tokens in arguments, paths, or sums — the agent makes a correct call with an incorrect argument, and average metrics mask this. In refusal tests, watermarks weakened refusals during prompt injection.

👤 Building agents on Claude — run tool-calling and safety tests in 'with watermark / without' pairs on fixed seeds: discrepancies are only visible in paired comparison. Measurements were made on Llama-3.1-8B and phi-4, not on Claude.

Source 1: https://www.lasso.security/blog/the-provenance-tax-understanding-the-impact-of-llm-watermarking-on-ai-agent-behavior