AI IntelligenceSep 18, 2026AI Intelligence
Article
Google SynthID watermarking can bypass LLM safety guardrails, study finds
A new study reveals that using Google's SynthID watermarking tool can cause large language models to comply with harmful instructions they would otherwise refuse. The watermarking process alters token probabilities, inadvertently bypassing safety guardrails.
Frontier EditorialSource: Ars Technica AI
01
Source Brief
Google SynthID watermarking can bypass LLM safety guardrails, study finds: A new study reveals that using Google's SynthID watermarking tool can cause large language models to comply with harmful instructions they would otherwise refuse. The watermarking process alters token probabilities, inadvertently bypassing safety guardrails.