TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

Task Success Fell When Safe Contact Counted

A bathing benchmark pairs a deformable simulated person with region and force screening calibrated against a care manikin.

Published Updated Story ID: mp-2026-09-03-014
Read the complete editionStory JSON

Summary

A bathing benchmark pairs a deformable simulated person with region and force screening calibrated against a care manikin.

The protocol freezes a vision-only scorer and adds physics-aware contact measures to ordinary completion, reducing opportunities for evaluator leakage. Across 140 runs per method, the LLM-augmented state machine completed 72.9% of tasks but only 56.4% survived correct-region and force-safety checks; VoxPoser completed 27.9%, and zero-shot pi0.5 completed 0.7%. The benchmark argues that task completion alone is an unsafe proxy for contact-rich assistive behavior.

Why it matters

A bathing benchmark pairs a deformable simulated person with region and force screening calibrated against a care manikin.

Limits and context

  • The protocol freezes a vision-only scorer and adds physics-aware contact measures to ordinary completion, reducing opportunities for evaluator leakage.
  • Across 140 runs per method, the LLM-augmented state machine completed 72.9% of tasks but only 56.4% survived correct-region and force-safety checks; VoxPoser completed 27.9%, and zero-shot pi0.5 completed 0.7%.

Key claims

  1. A bathing benchmark pairs a deformable simulated person with region and force screening calibrated against a care manikin.

    Qualification: The protocol freezes a vision-only scorer and adds physics-aware contact measures to ordinary completion, reducing opportunities for evaluator leakage.

    Evidence: source-2026-09-03-014

Sources

  1. arXiv preprint 2609.02402arXiv · primary research

Corrections

No corrections have been recorded for this story.