TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

robotics

The Robot Policy Gave Credit to Each Denoising Step

DIA learned values for partially denoised actions instead of assigning one environment-level reward to an entire action chunk.

Published Updated Story ID: mp-2026-09-14-011
Read the complete editionStory JSON

Summary

DIA learned values for partially denoised actions instead of assigning one environment-level reward to an entire action chunk.

Diffusion robot policies generate an action through multiple denoising steps, but common reinforcement-learning methods assign the same final credit to every intermediate choice. Denoising Intermediate Advantage learns a value function over partially denoised actions and combines that inner-step signal with the ordinary PPO advantage from the environment. Across Robomimic, FurnitureBench, Franka Kitchen and D3IL, the authors report higher final performance than the tested diffusion-policy fine-tuning methods, earlier arrival at successful states and greater movement away from the behavior-cloned starting distribution. The abstract does not supply one aggregate gain or physical-robot guarantee.

Why it matters

DIA learned values for partially denoised actions instead of assigning one environment-level reward to an entire action chunk.

Limits and context

  • The abstract does not supply one aggregate gain or physical-robot guarantee.

Key claims

  1. DIA learned values for partially denoised actions instead of assigning one environment-level reward to an entire action chunk.

    Qualification: The abstract does not supply one aggregate gain or physical-robot guarantee.

    Evidence: source-2026-09-14-011

Sources

  1. arXiv preprint 2609.12245arXiv · primary research

Corrections

No corrections have been recorded for this story.