TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

robotics

The Humanoid Learned to Walk to the Grasp

Reinforcement learning pushed a cloned policy past its synthetic-data ceiling and transferred a reach-and-pick skill to a real humanoid.

Published Updated Story ID: mp-2026-08-19-003
Read the complete editionStory JSON

Summary

Reinforcement learning pushed a cloned policy past its synthetic-data ceiling and transferred a reach-and-pick skill to a real humanoid.

FetchMan begins with synthetic demonstrations across more than 150,000 scenes, then refines the cloned policy with Flow-GRPO and a sparse reward. The authors report 73.3 percent zero-shot success when a Unitree G1 walked toward and grasped a single target in unseen real scenes. The result supports the sim-to-real recipe for one reach-and-pick policy; it does not establish general household manipulation.

Why it matters

Reinforcement learning pushed a cloned policy past its synthetic-data ceiling and transferred a reach-and-pick skill to a real humanoid.

Limits and context

  • The result supports the sim-to-real recipe for one reach-and-pick policy; it does not establish general household manipulation.

Key claims

  1. Reinforcement learning pushed a cloned policy past its synthetic-data ceiling and transferred a reach-and-pick skill to a real humanoid.

    Qualification: The result supports the sim-to-real recipe for one reach-and-pick policy; it does not establish general household manipulation.

    Evidence: source-2026-08-19-003

Sources

  1. arXiv preprint 2608.17027arXiv · primary research

Corrections

No corrections have been recorded for this story.