Adversarial Prompts
Topic archive • 3 matches
2026-09-18
Technology
Google SynthID watermarking can bypass LLM safety guardrails, study finds: A new study reveals that using Google's SynthID watermarking tool can cause large language models to comply with harmful instructions they would otherwise refuse. The watermarking process alters token probabilities, inadvertently bypassing safety guardrails.
Security • Ars Technica AI
Permalink
2026-09-15
Tips
Prompt Engineering
Tuning prompts against small evaluation sets risks model overfitting
Permalink
2026-09-14
Tips
Large Language Models
Audit identifies 12 data leaks and compliance risks in agentic LLM pipelines
Permalink