robotics
The Humanoid Learned to Walk to the Grasp
Reinforcement learning pushed a cloned policy past its synthetic-data ceiling and transferred a reach-and-pick skill to a real humanoid.

Summary
Reinforcement learning pushed a cloned policy past its synthetic-data ceiling and transferred a reach-and-pick skill to a real humanoid.
FetchMan begins with synthetic demonstrations across more than 150,000 scenes, then refines the cloned policy with Flow-GRPO and a sparse reward. The authors report 73.3 percent zero-shot success when a Unitree G1 walked toward and grasped a single target in unseen real scenes. The result supports the sim-to-real recipe for one reach-and-pick policy; it does not establish general household manipulation.
Why it matters
Reinforcement learning pushed a cloned policy past its synthetic-data ceiling and transferred a reach-and-pick skill to a real humanoid.
Limits and context
- The result supports the sim-to-real recipe for one reach-and-pick policy; it does not establish general household manipulation.
Key claims
Reinforcement learning pushed a cloned policy past its synthetic-data ceiling and transferred a reach-and-pick skill to a real humanoid.
Qualification: The result supports the sim-to-real recipe for one reach-and-pick policy; it does not establish general household manipulation.
Evidence: source-2026-08-19-003
Sources
- arXiv preprint 2608.17027arXiv · primary research
Corrections
No corrections have been recorded for this story.