Skip to content

Adversarial Prompts

Topic archive1 matches

Back to homeGEO summary endpoint

2026-09-18

Technology

  • Google SynthID watermarking can bypass LLM safety guardrails, study finds: A new study reveals that using Google's SynthID watermarking tool can cause large language models to comply with harmful instructions they would otherwise refuse. The watermarking process alters token probabilities, inadvertently bypassing safety guardrails.

    SecurityArs Technica AI

    Permalink