TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

Robot Success Hid Weak Instruction Following

Scenes with only one plausible task let embodied policies ignore language.

Published Updated Story ID: mp-2026-09-23-012
Read the complete editionStory JSON

Summary

Scenes with only one plausible task let embodied policies ignore language.

RoboFollow raises scene entropy by placing multiple valid task branches in one scene, so a policy must distinguish instructions rather than infer the sole available action. Nine evaluated VLA and world-action models that performed well in the easiest setting did not reliably transfer across progressively changed layouts and semantics after fine-tuning. Stronger vision-language backbones and several training adjustments did not close the gap. This is a diagnostic benchmark finding, not a claim that every deployed robot ignores language.

Why it matters

Scenes with only one plausible task let embodied policies ignore language.

Limits and context

  • Nine evaluated VLA and world-action models that performed well in the easiest setting did not reliably transfer across progressively changed layouts and semantics after fine-tuning.
  • Stronger vision-language backbones and several training adjustments did not close the gap.
  • This is a diagnostic benchmark finding, not a claim that every deployed robot ignores language.

Key claims

  1. Scenes with only one plausible task let embodied policies ignore language.

    Qualification: Nine evaluated VLA and world-action models that performed well in the easiest setting did not reliably transfer across progressively changed layouts and semantics after fine-tuning.

    Evidence: source-2026-09-23-012

Sources

  1. arXiv preprint 2609.25636arXiv · primary research

Corrections

No corrections have been recorded for this story.