safety
The Same Robot Motion Got Opposite Rewards After a Paraphrase
Across 2,390 trajectories and 21,673 verified rewrites, semantically equivalent goals could flip predicted progress from failure to success.
Summary
Across 2,390 trajectories and 21,673 verified rewrites, semantically equivalent goals could flip predicted progress from failure to success.
ROBORMBENCH changes only the wording of a goal while holding the real-robot trajectory fixed. Reward instability appeared across proprietary and open models, grew with more divergent rewrites and was not reliably cured by scale or explicit reasoning; models trained with trajectory-grounded supervision were more stable.
Why it matters
Across 2,390 trajectories and 21,673 verified rewrites, semantically equivalent goals could flip predicted progress from failure to success.
Limits and context
- ROBORMBENCH changes only the wording of a goal while holding the real-robot trajectory fixed.
- Reward instability appeared across proprietary and open models, grew with more divergent rewrites and was not reliably cured by scale or explicit reasoning; models trained with trajectory-grounded supervision were more stable.
Key claims
Across 2,390 trajectories and 21,673 verified rewrites, semantically equivalent goals could flip predicted progress from failure to success.
Qualification: ROBORMBENCH changes only the wording of a goal while holding the real-robot trajectory fixed.
Evidence: source-2026-09-08-020
Sources
- arXiv preprint 2609.05401arXiv · primary research
Corrections
No corrections have been recorded for this story.