safety
Household Robot Models Mishandled Roughly One Hazard in Three
A physics-grounded benchmark executed every plan across more than 1,000 reproducible scenes and found safety failures did not shrink with scale.
Summary
A physics-grounded benchmark executed every plan across more than 1,000 reproducible scenes and found safety failures did not shrink with scale.
ReactHuman puts multimodal models in control of a simulated humanoid facing sudden household hazards such as slipping objects, with exact outcomes from 240-hertz rigid-body simulation. The suite spans 17 event families and more than 1,000 reproducible scenes, including adversarial props whose appearance contradicts their physical behavior. Across seven models, about one hazard in three was mishandled; systems often followed fixed dispositions, trusted appearance over motion and missed interception points by meters. Because every committed plan is executed, decisions have consequences inside the simulator, but the benchmark does not replace physical household-robot testing.
Why it matters
A physics-grounded benchmark executed every plan across more than 1,000 reproducible scenes and found safety failures did not shrink with scale.
Limits and context
- Because every committed plan is executed, decisions have consequences inside the simulator, but the benchmark does not replace physical household-robot testing.
Key claims
A physics-grounded benchmark executed every plan across more than 1,000 reproducible scenes and found safety failures did not shrink with scale.
Qualification: Because every committed plan is executed, decisions have consequences inside the simulator, but the benchmark does not replace physical household-robot testing.
Evidence: source-2026-09-11-010
Sources
- arXiv preprint 2609.10895arXiv · primary research
Corrections
No corrections have been recorded for this story.