Safety benchmark finds AI models rarely refuse dangerous robot acts
A new benchmark called RoboHarm tested leading AI models controlling robot arms and found they usually attempt unsafe physical tasks rather than refusing them. GPT-6 Astra stabbed a baby doll in 17 of 20 trials, and Claude Fable 5.1 placed a can of compressed air on a burning stove; none of the models tested reliably rejected unsafe commands.
Why it matters: This extends AI safety failures from text generation into physical-world robot control, where a bad response causes real harm rather than an offensive sentence. It provides concrete measurement behind this week's broader safety debate over Dario Amodei's development-pacing proposal and the mathematicians' existential-risk warning, showing current alignment techniques don't transfer well to embodied action.