TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

robotics

The Human Corrected the Policy Without Running the Robot

HIL-UMI queried the current policy during handheld demonstrations and collected data where action disagreement was high.

Published Updated Story ID: mp-2026-09-18-008
Read the complete editionStory JSON

Summary

HIL-UMI queried the current policy during handheld demonstrations and collected data where action disagreement was high.

Interactive robot post-training usually means repeatedly running the policy on hardware and waiting for human intervention. HIL-UMI instead queries the current policy on the observation stream from a handheld Universal Manipulation Interface without executing the predicted motion. An energy score targets collection where human and policy trajectories disagree, while low advantage predictions identify segments that refine a progress estimator. Across four physical tasks, the policy improved over supervised fine-tuning and beat HG-DAgger on table cleanup with lower per-frame collection time.

Why it matters

HIL-UMI queried the current policy during handheld demonstrations and collected data where action disagreement was high.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. HIL-UMI queried the current policy during handheld demonstrations and collected data where action disagreement was high.

    Evidence: source-2026-09-18-008

Sources

  1. arXiv preprint 2609.20659arXiv · primary research

Corrections

No corrections have been recorded for this story.