robotics
The Human Corrected the Policy Without Running the Robot
HIL-UMI queried the current policy during handheld demonstrations and collected data where action disagreement was high.

Summary
HIL-UMI queried the current policy during handheld demonstrations and collected data where action disagreement was high.
Interactive robot post-training usually means repeatedly running the policy on hardware and waiting for human intervention. HIL-UMI instead queries the current policy on the observation stream from a handheld Universal Manipulation Interface without executing the predicted motion. An energy score targets collection where human and policy trajectories disagree, while low advantage predictions identify segments that refine a progress estimator. Across four physical tasks, the policy improved over supervised fine-tuning and beat HG-DAgger on table cleanup with lower per-frame collection time.
Why it matters
HIL-UMI queried the current policy during handheld demonstrations and collected data where action disagreement was high.
Limits and context
No additional limitation was separately recorded.
Key claims
HIL-UMI queried the current policy during handheld demonstrations and collected data where action disagreement was high.
Evidence: source-2026-09-18-008
Sources
- arXiv preprint 2609.20659arXiv · primary research
Corrections
No corrections have been recorded for this story.