TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

safety

The Same Robot Motion Got Opposite Rewards After a Paraphrase

Across 2,390 trajectories and 21,673 verified rewrites, semantically equivalent goals could flip predicted progress from failure to success.

Published Updated Story ID: mp-2026-09-08-018
Read the complete editionStory JSON

Summary

Across 2,390 trajectories and 21,673 verified rewrites, semantically equivalent goals could flip predicted progress from failure to success.

ROBORMBENCH changes only the wording of a goal while holding the real-robot trajectory fixed. Reward instability appeared across proprietary and open models, grew with more divergent rewrites and was not reliably cured by scale or explicit reasoning; models trained with trajectory-grounded supervision were more stable.

Why it matters

Across 2,390 trajectories and 21,673 verified rewrites, semantically equivalent goals could flip predicted progress from failure to success.

Limits and context

  • ROBORMBENCH changes only the wording of a goal while holding the real-robot trajectory fixed.
  • Reward instability appeared across proprietary and open models, grew with more divergent rewrites and was not reliably cured by scale or explicit reasoning; models trained with trajectory-grounded supervision were more stable.

Key claims

  1. Across 2,390 trajectories and 21,673 verified rewrites, semantically equivalent goals could flip predicted progress from failure to success.

    Qualification: ROBORMBENCH changes only the wording of a goal while holding the real-robot trajectory fixed.

    Evidence: source-2026-09-08-020

Sources

  1. arXiv preprint 2609.05401arXiv · primary research

Corrections

No corrections have been recorded for this story.