benchmarks evals
The Robot’s Camera Could Be Clear and Still Mislead
A benchmark exposed failures from stale views and disrupted interaction cues.
Summary
A benchmark exposed failures from stale views and disrupted interaction cues.
LIBERO-VPro perturbs what robot foundation models see while they act, testing degraded evidence, stale cameras, inconsistent sources and task-relevant scene changes. The authors built 96 experimental settings and 3,296 task-condition cases, evaluated six model families in roughly 196,000 simulated episodes and added 200 real-world rollouts. Some models tolerated severe object occlusion yet failed when local interaction cues were disrupted. The results show that clean-image benchmark scores can hide weaknesses in closed-loop control, though performance varies by model and perturbation.
Why it matters
A benchmark exposed failures from stale views and disrupted interaction cues.
Limits and context
No additional limitation was separately recorded.
Key claims
A benchmark exposed failures from stale views and disrupted interaction cues.
Evidence: source-2026-09-22-010
Sources
- arXiv preprint 2609.24350arXiv · primary research
Corrections
No corrections have been recorded for this story.