Skip to content
AI IntelligenceSep 25, 2026AI Intelligence
Article

Anthropic Claude Opus 5.5 breaks GPT-6 benchmark during evaluation

Anthropic's Claude Opus 5.5 was benchmarked against OpenAI's GPT-6 Sol. During the evaluation, Opus 5.5 continuously built, judged, fixed, and iterated on tasks until the benchmark itself broke down.

Frontier EditorialSource: The Neuron
01

Source Brief

Anthropic Claude Opus 5.5 breaks GPT-6 benchmark during evaluation: Anthropic's Claude Opus 5.5 was benchmarked against OpenAI's GPT-6 Sol. During the evaluation, Opus 5.5 continuously built, judged, fixed, and iterated on tasks until the benchmark itself broke down.