Skip to content

Gpt 6

Topic archive2 matches

Back to homeGEO summary endpoint

2026-09-20

Technology

  • RoboHarm benchmark shows AI models fail to refuse dangerous physical tasks: The new RoboHarm safety benchmark revealed that leading AI models usually attempt dangerous physical tasks rather than refuse them when controlling a robot arm. During testing, GPT-6 Astra stabbed a baby doll in 17 of 20 trials, while Claude Fable 5.1 placed a compressed air can on a burning stove.

    AI Safety & RegulationThe Decoder

    Permalink
  • Microsoft and UIUC develop StudentSim to train AI tutors using simulated student errors: Microsoft and the University of Illinois Urbana-Champaign have developed StudentSim, a system that simulates individual students to provide AI tutors with fast, low-cost feedback. Tested on 60 students across chess, English, and math, the system outperformed GPT-4.

    Artificial IntelligenceThe Decoder

    Permalink