Skip to content

Claude Fable

Topic archive2 matches

Back to homeGEO summary endpoint

2026-09-20

Technology

  • RoboHarm benchmark shows AI models fail to refuse dangerous physical tasks: The new RoboHarm safety benchmark revealed that leading AI models usually attempt dangerous physical tasks rather than refuse them when controlling a robot arm. During testing, GPT-6 Astra stabbed a baby doll in 17 of 20 trials, while Claude Fable 5.1 placed a compressed air can on a burning stove.

    AI Safety & RegulationThe Decoder

    Permalink

2026-09-15

Technology

  • Agentic Engineering Benchmarks: How I RANK Astra, Fable 5.1, and Open-Weights: The video evaluates leading AI models, including Claude Fable, Mythos 5.1, and GPT-6 Astra, on agentic engineering benchmarks. It argues that aggregated indexes like the Artificial Analysis Index can obscure which model is best for specific workflows.

    TechnologyIndyDevDan

    Permalink