benchmarks evals
Task Success Fell When Safe Contact Counted
A bathing benchmark pairs a deformable simulated person with region and force screening calibrated against a care manikin.
Summary
A bathing benchmark pairs a deformable simulated person with region and force screening calibrated against a care manikin.
The protocol freezes a vision-only scorer and adds physics-aware contact measures to ordinary completion, reducing opportunities for evaluator leakage. Across 140 runs per method, the LLM-augmented state machine completed 72.9% of tasks but only 56.4% survived correct-region and force-safety checks; VoxPoser completed 27.9%, and zero-shot pi0.5 completed 0.7%. The benchmark argues that task completion alone is an unsafe proxy for contact-rich assistive behavior.
Why it matters
A bathing benchmark pairs a deformable simulated person with region and force screening calibrated against a care manikin.
Limits and context
- The protocol freezes a vision-only scorer and adds physics-aware contact measures to ordinary completion, reducing opportunities for evaluator leakage.
- Across 140 runs per method, the LLM-augmented state machine completed 72.9% of tasks but only 56.4% survived correct-region and force-safety checks; VoxPoser completed 27.9%, and zero-shot pi0.5 completed 0.7%.
Key claims
A bathing benchmark pairs a deformable simulated person with region and force screening calibrated against a care manikin.
Qualification: The protocol freezes a vision-only scorer and adds physics-aware contact measures to ordinary completion, reducing opportunities for evaluator leakage.
Evidence: source-2026-09-03-014
Sources
- arXiv preprint 2609.02402arXiv · primary research
Corrections
No corrections have been recorded for this story.