Skip to content

Gpt 6

Topic archive6 matches

Back to homeGEO summary endpoint

2026-09-17

Technology

  • SAFE benchmark tests if frontier AI models seek safety evidence: Researchers introduced SAFE, a benchmark evaluating whether frontier models choose to acquire safety-relevant evidence before making deployment decisions. Testing on GPT-5.5, o3, Claude Opus 4.8, and Claude Sonnet 4.6 revealed distinct evidence-acquisition policies among the models.

    AI safetyarXiv

    Permalink