robotics
A Humanoid Learned Motion Straight From Monocular Video
BeyondRetarget maps RGB video directly to robot motion instead of first reconstructing a human skeleton.

Summary
BeyondRetarget maps RGB video directly to robot motion instead of first reconstructing a human skeleton.
BeyondRetarget learns a robot-oriented representation directly from monocular RGB video, avoiding a separate human-motion reconstruction and retargeting stage. A contact-aware optimization step improves temporal consistency and physical plausibility. The authors report higher motion accuracy, execution success and robustness with lower latency in simulation and on real humanoid robots, but the abstract does not supply absolute success rates or establish performance outside the evaluated motions and platforms.
Why it matters
BeyondRetarget maps RGB video directly to robot motion instead of first reconstructing a human skeleton.
Limits and context
- The authors report higher motion accuracy, execution success and robustness with lower latency in simulation and on real humanoid robots, but the abstract does not supply absolute success rates or establish performance outside the evaluated motions and platforms.
Key claims
BeyondRetarget maps RGB video directly to robot motion instead of first reconstructing a human skeleton.
Qualification: The authors report higher motion accuracy, execution success and robustness with lower latency in simulation and on real humanoid robots, but the abstract does not supply absolute success rates or establish performance outside the evaluated motions and platforms.
Evidence: source-2026-09-25-013
Sources
- arXiv preprint 2609.29850arXiv · primary research
Corrections
No corrections have been recorded for this story.