robotics
The Robot Video Looked Right Until It Had to Act
Eleven world models struggled to turn egocentric human demonstrations into robot videos with consistent embodiments, contacts and completed tasks.
Summary
Eleven world models struggled to turn egocentric human demonstrations into robot videos with consistent embodiments, contacts and completed tasks.
H2R-Bench pairs a human demonstration with target robot constraints and source-grounded annotations for goals, action events, functional contacts and object responses. Across six manipulation families and two robot embodiments, even leading video generators often failed embodiment consistency, functional interaction or task execution. The benchmark separates those failures from general video quality, exposing why visually plausible clips are not automatically useful robot-training data. It evaluates generated videos; it does not show that the models safely control physical robots.
Why it matters
Eleven world models struggled to turn egocentric human demonstrations into robot videos with consistent embodiments, contacts and completed tasks.
Limits and context
- The benchmark separates those failures from general video quality, exposing why visually plausible clips are not automatically useful robot-training data.
- It evaluates generated videos; it does not show that the models safely control physical robots.
Key claims
Eleven world models struggled to turn egocentric human demonstrations into robot videos with consistent embodiments, contacts and completed tasks.
Qualification: The benchmark separates those failures from general video quality, exposing why visually plausible clips are not automatically useful robot-training data.
Evidence: source-2026-08-14-006
Sources
- arXiv preprint 2608.13049arXiv · primary research
Corrections
No corrections have been recorded for this story.