robotics
A Robot Could Finish and Still Leave a Hazard
SafeStage checks risks before, during and after manipulation across 97 scenarios.

Summary
SafeStage checks risks before, during and after manipulation across 97 scenarios.
SafeStage separates task success from manipulation safety in 97 designed scenarios. It evaluates initial hazards, unsafe contacts and trajectories during execution, and unstable or unsafe final states after nominal completion. Tests of representative vision-language-action and world-model-based robot policies found that completed tasks could still contain safety violations, with different policies failing at different stages. The benchmark diagnoses failure modes in a controlled protocol; it does not certify a real robot as safe.
Why it matters
SafeStage checks risks before, during and after manipulation across 97 scenarios.
Limits and context
- The benchmark diagnoses failure modes in a controlled protocol; it does not certify a real robot as safe.
Key claims
SafeStage checks risks before, during and after manipulation across 97 scenarios.
Qualification: The benchmark diagnoses failure modes in a controlled protocol; it does not certify a real robot as safe.
Evidence: source-2026-09-21-008
Sources
- arXiv preprint 2609.21223arXiv · primary research
Corrections
No corrections have been recorded for this story.