robotics
Six Hundred Thirty Thousand Human Videos Learned a Robot’s Shape
Robotizing observations and actions raised real-world completion from 51.5% to 80.3% as pretraining scale increased.
Summary
Robotizing observations and actions raised real-world completion from 51.5% to 80.3% as pretraining scale increased.
HuRo converts heterogeneous human videos into robot-aligned observations and action trajectories, inferring missing intermediate signals across annotation levels. The resulting dataset contains roughly 630,000 episodes and 142 million processed frames from five video sources. Across four real manipulation tasks, scaling the robotized pretraining data increased overall completion from 51.5% to 80.3%; out-of-distribution completion under spatial and visual shifts rose from 34.9% to 72.2%. Ablations attribute some robustness to visual robotization and favor end-to-end retargeted actions over visual-only transfer. The evidence covers the authors’ four tasks, not arbitrary robots or internet video.
Why it matters
Robotizing observations and actions raised real-world completion from 51.5% to 80.3% as pretraining scale increased.
Limits and context
- Ablations attribute some robustness to visual robotization and favor end-to-end retargeted actions over visual-only transfer.
- The evidence covers the authors’ four tasks, not arbitrary robots or internet video.
Key claims
Robotizing observations and actions raised real-world completion from 51.5% to 80.3% as pretraining scale increased.
Qualification: Ablations attribute some robustness to visual robotization and favor end-to-end retargeted actions over visual-only transfer.
Evidence: source-2026-09-11-009
Sources
- arXiv preprint 2609.10706arXiv · primary research
Corrections
No corrections have been recorded for this story.