TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

robotics

The Robot Practiced Only Its Weak Steps

Targeted real-world reinforcement learning lifted two long-task success rates with limited rollouts.

Published Updated Story ID: mp-2026-09-21-010
Read the complete editionStory JSON

Summary

Targeted real-world reinforcement learning lifted two long-task success rates with limited rollouts.

PARTS keeps a pretrained robot policy frozen for routine portions of a task and trains residual corrections at bottleneck subtasks. Agent-generated selectors and success checks provide local rewards even when the complete task rarely succeeds. On bimanual YAM and single-arm Franka experiments, full-task success rose from 32% to 61% and 50% to 95%, respectively, with tens of minutes of real-world reinforcement-learning rollouts per task on average. Humans still chose bottlenecks and handled physical resets.

Why it matters

Targeted real-world reinforcement learning lifted two long-task success rates with limited rollouts.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. Targeted real-world reinforcement learning lifted two long-task success rates with limited rollouts.

    Evidence: source-2026-09-21-010

Sources

  1. arXiv preprint 2609.21788arXiv · primary research

Corrections

No corrections have been recorded for this story.