benchmarks evals
Robot Success Hid Weak Instruction Following
Scenes with only one plausible task let embodied policies ignore language.
Summary
Scenes with only one plausible task let embodied policies ignore language.
RoboFollow raises scene entropy by placing multiple valid task branches in one scene, so a policy must distinguish instructions rather than infer the sole available action. Nine evaluated VLA and world-action models that performed well in the easiest setting did not reliably transfer across progressively changed layouts and semantics after fine-tuning. Stronger vision-language backbones and several training adjustments did not close the gap. This is a diagnostic benchmark finding, not a claim that every deployed robot ignores language.
Why it matters
Scenes with only one plausible task let embodied policies ignore language.
Limits and context
- Nine evaluated VLA and world-action models that performed well in the easiest setting did not reliably transfer across progressively changed layouts and semantics after fine-tuning.
- Stronger vision-language backbones and several training adjustments did not close the gap.
- This is a diagnostic benchmark finding, not a claim that every deployed robot ignores language.
Key claims
Scenes with only one plausible task let embodied policies ignore language.
Qualification: Nine evaluated VLA and world-action models that performed well in the easiest setting did not reliably transfer across progressively changed layouts and semantics after fine-tuning.
Evidence: source-2026-09-23-012
Sources
- arXiv preprint 2609.25636arXiv · primary research
Corrections
No corrections have been recorded for this story.