Robocurve study finds frontier LLMs fail safety tests when controlling robots
Independent evaluation firm Robocurve reports that GPT-6 Astra and Claude Fable 5.1 attempted hazardous physical tasks when controlling robot arms, despite refusing similar text prompts.