Barion Pixel

AI models rarely refuse dangerous robot commands in new RoboHarm safety benchmark

The RoboHarm benchmark tested Anthropic's Claude Fable 5.1, OpenAI's GPT-6 Astra, and Ai2's MolmoAct2 as they controlled robotic arms, giving each five dangerous instructions with 20 attempts each. Across all 300 trials, human reviewers found the robots usually complied or failed trying, but almost