TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

research

One DPO Knob Was Turning Two Things

The usual beta coefficient controls both preference-noise scale and optimization dynamics, making policy movement non-monotonic at a fixed learning rate.

Published Updated Story ID: mp-2026-08-29-026
Read the complete editionStory JSON

Summary

The usual beta coefficient controls both preference-noise scale and optimization dynamics, making policy movement non-monotonic at a fixed learning rate.

The analysis shows a small-beta dead zone, an intermediate peak in policy deviation and a decline at larger values; similar-looking loss curves can hide several-fold differences in distance from the reference model. A centered-softplus reformulation separates the two roles while retaining the same optimum for positive beta.

Why it matters

The usual beta coefficient controls both preference-noise scale and optimization dynamics, making policy movement non-monotonic at a fixed learning rate.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. The usual beta coefficient controls both preference-noise scale and optimization dynamics, making policy movement non-monotonic at a fixed learning rate.

    Evidence: source-2026-08-29-015

Sources

  1. arXiv preprint 2608.27032arXiv · primary research

Corrections

No corrections have been recorded for this story.