robotics
The Robot Policy Gave Credit to Each Denoising Step
DIA learned values for partially denoised actions instead of assigning one environment-level reward to an entire action chunk.
Summary
DIA learned values for partially denoised actions instead of assigning one environment-level reward to an entire action chunk.
Diffusion robot policies generate an action through multiple denoising steps, but common reinforcement-learning methods assign the same final credit to every intermediate choice. Denoising Intermediate Advantage learns a value function over partially denoised actions and combines that inner-step signal with the ordinary PPO advantage from the environment. Across Robomimic, FurnitureBench, Franka Kitchen and D3IL, the authors report higher final performance than the tested diffusion-policy fine-tuning methods, earlier arrival at successful states and greater movement away from the behavior-cloned starting distribution. The abstract does not supply one aggregate gain or physical-robot guarantee.
Why it matters
DIA learned values for partially denoised actions instead of assigning one environment-level reward to an entire action chunk.
Limits and context
- The abstract does not supply one aggregate gain or physical-robot guarantee.
Key claims
DIA learned values for partially denoised actions instead of assigning one environment-level reward to an entire action chunk.
Qualification: The abstract does not supply one aggregate gain or physical-robot guarantee.
Evidence: source-2026-09-14-011
Sources
- arXiv preprint 2609.12245arXiv · primary research
Corrections
No corrections have been recorded for this story.