TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Robot’s Camera Could Be Clear and Still Mislead

A benchmark exposed failures from stale views and disrupted interaction cues.

Published Updated Story ID: mp-2026-09-22-010
Read the complete editionStory JSON

Summary

A benchmark exposed failures from stale views and disrupted interaction cues.

LIBERO-VPro perturbs what robot foundation models see while they act, testing degraded evidence, stale cameras, inconsistent sources and task-relevant scene changes. The authors built 96 experimental settings and 3,296 task-condition cases, evaluated six model families in roughly 196,000 simulated episodes and added 200 real-world rollouts. Some models tolerated severe object occlusion yet failed when local interaction cues were disrupted. The results show that clean-image benchmark scores can hide weaknesses in closed-loop control, though performance varies by model and perturbation.

Why it matters

A benchmark exposed failures from stale views and disrupted interaction cues.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. A benchmark exposed failures from stale views and disrupted interaction cues.

    Evidence: source-2026-09-22-010

Sources

  1. arXiv preprint 2609.24350arXiv · primary research

Corrections

No corrections have been recorded for this story.