benchmarks evals
One Word Changed the Robot Task—and Several Models Lost the Constraints
A real-robot benchmark separates instructions that should preserve an action from those that should change it, then adds multiple constraints.
Summary
A real-robot benchmark separates instructions that should preserve an action from those that should change it, then adds multiple constraints.
One Word, Different Action builds paired instructions around physical decision states and executable actions. Some wording changes leave the task intact and test decision invariance; others change the task and test decision sensitivity. Modern models approached saturation when a single constraint changed, according to the authors, but several degraded when one action decision had to integrate multiple requirements. The benchmark suggests the remaining problem is less about noticing an isolated word and more about composing constraints reliably under real visual grounding.
Why it matters
A real-robot benchmark separates instructions that should preserve an action from those that should change it, then adds multiple constraints.
Limits and context
- Modern models approached saturation when a single constraint changed, according to the authors, but several degraded when one action decision had to integrate multiple requirements.
Key claims
A real-robot benchmark separates instructions that should preserve an action from those that should change it, then adds multiple constraints.
Qualification: Modern models approached saturation when a single constraint changed, according to the authors, but several degraded when one action decision had to integrate multiple requirements.
Evidence: source-2026-09-08-014
Sources
- arXiv preprint 2609.05260arXiv · primary research
Corrections
No corrections have been recorded for this story.