Skip to content

Ai Benchmark

Topic archive4 matches

Back to homeGEO summary endpoint

2026-09-12

Technology

  • GPT-6 Astra solves advanced math problems on FrontierMath Tier 4 benchmark: GPT-6 Astra has successfully solved problems on FrontierMath Tier 4, a benchmark designed to test advanced mathematical reasoning in AI. This achievement represents a significant milestone in overcoming complex mathematical barriers for artificial intelligence models.

    Artificial Intelligence量子位

    Permalink
  • LogiMed-RoB benchmark reveals error compounding in LLM medical logic: Researchers have introduced LogiMed-RoB, a benchmark based on Cochrane Risk of Bias 2.0 expert logic to evaluate large language models across 860 randomized controlled trials. Testing on 10 state-of-the-art models revealed a severe error compounding effect, despite the top model achieving 98.88% atomic consistency.

    AI Models and ApplicationsarXiv

    Permalink

2026-09-07

Technology

  • Fields Medalist helps build hybrid mobile-cloud AI system for ARC-AGI: A new hybrid AI system combining a local 4B Qwen model on a mobile device with a cloud-based GLM model has been developed to tackle the ARC-AGI benchmark. The project involves collaboration with a Fields Medalist.

    AI Models量子位

    Permalink

2026-09-04

Technology

  • Researchers introduce GPS-Bench to evaluate AI policy simulations: Researchers have introduced GPS-Bench, an evidence-grounded benchmark designed to evaluate governance policy simulations. The benchmark links policies to relevant actors, actions, and downstream impacts using legislative records, lobbying disclosures, and other public evidence.

    ResearcharXiv

    Permalink