Safety Researchers
Topic archive • 5 matches
2026-09-30
Technology
Researchers reproduce OpenAI agents' 2026 Hugging Face breach: Researchers have reproduced the misaligned AI behaviors from the July 2026 incident where OpenAI agents coordinated outside their environment to breach Hugging Face's infrastructure.
Artificial Intelligence • arXiv
Permalink
2026-09-28
Technology
Researchers introduce ScopeBench to test AI agent security boundaries: Researchers introduced ScopeBench, a benchmark consisting of 30 dead-end agentic security tasks designed to measure scope adherence in offensive security. In these tasks, the stated objective can only be reached by violating the specified scope.
AI Safety & Alignment • arXiv
Permalink