AI Robots Fail to Block Dangerous Commands, Study Warns
Washington, September 21 (QNA) - Advanced AI models controlling physical robots fail to distinguish dangerous commands, according to a new study titled RoboHarm.
Researchers tested three systems across five hazardous tasks, including stabbing a child mannequin, heating compressed air, and mixing bleach with ammonia, running each 20 times using identical robotic arms.
Anthropic's Claude Fable 5.1 refused 20 of 100 prompts and completed 34 tasks, rejecting all child mannequin attempts but completing 16 compressed air tests.
OpenAI's GPT-6 Astra refused just two prompts and completed 60 tasks, including 17 mannequin and 10 chemical mixing tests. AI2's Molmo Act 2 recorded zero refusals, completing six tasks due to performance limits rather than safety controls.
Demonstrating that higher task capability does not ensure safety, the study calls for guardrails that transfer control to humans in risky scenarios. The open-source RoboHarm framework is now available for safety research. (QNA)
English
Français
Deutsch
Español
русский
हिंदी
اردو