TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

robotics

Six Hundred Thirty Thousand Human Videos Learned a Robot’s Shape

Robotizing observations and actions raised real-world completion from 51.5% to 80.3% as pretraining scale increased.

Published Updated Story ID: mp-2026-09-11-009
Read the complete editionStory JSON

Summary

Robotizing observations and actions raised real-world completion from 51.5% to 80.3% as pretraining scale increased.

HuRo converts heterogeneous human videos into robot-aligned observations and action trajectories, inferring missing intermediate signals across annotation levels. The resulting dataset contains roughly 630,000 episodes and 142 million processed frames from five video sources. Across four real manipulation tasks, scaling the robotized pretraining data increased overall completion from 51.5% to 80.3%; out-of-distribution completion under spatial and visual shifts rose from 34.9% to 72.2%. Ablations attribute some robustness to visual robotization and favor end-to-end retargeted actions over visual-only transfer. The evidence covers the authors’ four tasks, not arbitrary robots or internet video.

Why it matters

Robotizing observations and actions raised real-world completion from 51.5% to 80.3% as pretraining scale increased.

Limits and context

  • Ablations attribute some robustness to visual robotization and favor end-to-end retargeted actions over visual-only transfer.
  • The evidence covers the authors’ four tasks, not arbitrary robots or internet video.

Key claims

  1. Robotizing observations and actions raised real-world completion from 51.5% to 80.3% as pretraining scale increased.

    Qualification: Ablations attribute some robustness to visual robotization and favor end-to-end retargeted actions over visual-only transfer.

    Evidence: source-2026-09-11-009

Sources

  1. arXiv preprint 2609.10706arXiv · primary research

Corrections

No corrections have been recorded for this story.