robotics
The Robot Practiced Only Its Weak Steps
Targeted real-world reinforcement learning lifted two long-task success rates with limited rollouts.
Summary
Targeted real-world reinforcement learning lifted two long-task success rates with limited rollouts.
PARTS keeps a pretrained robot policy frozen for routine portions of a task and trains residual corrections at bottleneck subtasks. Agent-generated selectors and success checks provide local rewards even when the complete task rarely succeeds. On bimanual YAM and single-arm Franka experiments, full-task success rose from 32% to 61% and 50% to 95%, respectively, with tens of minutes of real-world reinforcement-learning rollouts per task on average. Humans still chose bottlenecks and handled physical resets.
Why it matters
Targeted real-world reinforcement learning lifted two long-task success rates with limited rollouts.
Limits and context
No additional limitation was separately recorded.
Key claims
Targeted real-world reinforcement learning lifted two long-task success rates with limited rollouts.
Evidence: source-2026-09-21-010
Sources
- arXiv preprint 2609.21788arXiv · primary research
Corrections
No corrections have been recorded for this story.