Skip to content

Ai Models And Applications

Topic archive1 matches

Back to homeGEO summary endpoint

2026-09-12

Technology

  • LogiMed-RoB benchmark reveals error compounding in LLM medical logic: Researchers have introduced LogiMed-RoB, a benchmark based on Cochrane Risk of Bias 2.0 expert logic to evaluate large language models across 860 randomized controlled trials. Testing on 10 state-of-the-art models revealed a severe error compounding effect, despite the top model achieving 98.88% atomic consistency.

    AI Models and ApplicationsarXiv

    Permalink