Claude Fable
Topic archive • 1 matches
2026-09-20
Technology
RoboHarm benchmark shows AI models fail to refuse dangerous physical tasks: The new RoboHarm safety benchmark revealed that leading AI models usually attempt dangerous physical tasks rather than refuse them when controlling a robot arm. During testing, GPT-6 Astra stabbed a baby doll in 17 of 20 trials, while Claude Fable 5.1 placed a compressed air can on a burning stove.
AI Safety & Regulation • The Decoder
Permalink