TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

robotics

A Robot Estimated Its Own Training Noise Offline

Policy-Calibrated DAgger matched the best hindsight noise level without sweeping candidates.

Published Updated Story ID: mp-2026-09-28-026
Read the complete editionStory JSON

Summary

Policy-Calibrated DAgger matched the best hindsight noise level without sweeping candidates.

Policy-Calibrated DAgger samples a diffusion policy's action distribution along expert trajectories and relates that spread to closed-loop error. Partial denoising guides evaluation near a recorded path, while a correction estimates the unguided error used to choose data-collection noise. In a cluttered engine-lever simulator and a planar reacher, the resulting policies beat aggregation without noise and matched the best noise level selected after a sweep. The evidence is simulation-only and limited to the tested generative policy and tasks.

Why it matters

Policy-Calibrated DAgger matched the best hindsight noise level without sweeping candidates.

Limits and context

  • The evidence is simulation-only and limited to the tested generative policy and tasks.

Key claims

  1. Policy-Calibrated DAgger matched the best hindsight noise level without sweeping candidates.

    Qualification: The evidence is simulation-only and limited to the tested generative policy and tasks.

    Evidence: source-2026-09-28-015

Sources

  1. arXiv preprint 2609.30462arXiv · primary research

Corrections

No corrections have been recorded for this story.