robotics
A Robot Estimated Its Own Training Noise Offline
Policy-Calibrated DAgger matched the best hindsight noise level without sweeping candidates.
Summary
Policy-Calibrated DAgger matched the best hindsight noise level without sweeping candidates.
Policy-Calibrated DAgger samples a diffusion policy's action distribution along expert trajectories and relates that spread to closed-loop error. Partial denoising guides evaluation near a recorded path, while a correction estimates the unguided error used to choose data-collection noise. In a cluttered engine-lever simulator and a planar reacher, the resulting policies beat aggregation without noise and matched the best noise level selected after a sweep. The evidence is simulation-only and limited to the tested generative policy and tasks.
Why it matters
Policy-Calibrated DAgger matched the best hindsight noise level without sweeping candidates.
Limits and context
- The evidence is simulation-only and limited to the tested generative policy and tasks.
Key claims
Policy-Calibrated DAgger matched the best hindsight noise level without sweeping candidates.
Qualification: The evidence is simulation-only and limited to the tested generative policy and tasks.
Evidence: source-2026-09-28-015
Sources
- arXiv preprint 2609.30462arXiv · primary research
Corrections
No corrections have been recorded for this story.