Ai Benchmark
Topic archive • 2 matches
2026-09-12
Technology
GPT-6 Astra solves advanced math problems on FrontierMath Tier 4 benchmark: GPT-6 Astra has successfully solved problems on FrontierMath Tier 4, a benchmark designed to test advanced mathematical reasoning in AI. This achievement represents a significant milestone in overcoming complex mathematical barriers for artificial intelligence models.
Artificial Intelligence • 量子位
PermalinkLogiMed-RoB benchmark reveals error compounding in LLM medical logic: Researchers have introduced LogiMed-RoB, a benchmark based on Cochrane Risk of Bias 2.0 expert logic to evaluate large language models across 860 randomized controlled trials. Testing on 10 state-of-the-art models revealed a severe error compounding effect, despite the top model achieving 98.88% atomic consistency.
AI Models and Applications • arXiv
Permalink