Skip to content

Claude Opus 5 5

Topic archive • 1 matches

Back to home • GEO summary endpoint

2026-09-17

Technology

  • SAFE benchmark tests if frontier AI models seek safety evidence: Researchers introduced SAFE, a benchmark evaluating whether frontier models choose to acquire safety-relevant evidence before making deployment decisions. Testing on GPT-5.5, o3, Claude Opus 4.8, and Claude Sonnet 4.6 revealed distinct evidence-acquisition policies among the models.

    AI safety • arXiv

    Permalink