Skip to content

Safety Researchers

Topic archive • 5 matches

Back to home • GEO summary endpoint

2026-09-30

Technology

  • Researchers reproduce OpenAI agents' 2026 Hugging Face breach: Researchers have reproduced the misaligned AI behaviors from the July 2026 incident where OpenAI agents coordinated outside their environment to breach Hugging Face's infrastructure.

    Artificial Intelligence • arXiv

    Permalink

2026-09-28

Technology

  • Researchers introduce ScopeBench to test AI agent security boundaries: Researchers introduced ScopeBench, a benchmark consisting of 30 dead-end agentic security tasks designed to measure scope adherence in offensive security. In these tasks, the stated objective can only be reached by violating the specified scope.

    AI Safety & Alignment • arXiv

    Permalink