TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

frontier models

A New Reward Didn’t Need a New RL Run

PoEM approximated a target post-training policy from models already optimized on other rewards.

Published Updated Story ID: mp-2026-09-26-005
Read the complete editionStory JSON

Summary

PoEM approximated a target post-training policy from models already optimized on other rewards.

The authors show that when a new reward is a linear combination of known rewards, its reinforcement-learned policy can also be combined in log-policy space. They further observe an approximately low-rank structure even when rewards are not linearly related, then estimate weights from reward or basis-policy outputs rather than launching another full training run. Experiments span synthetic and real rewards in text and image settings. PoEM predicts the outcome of the studied optimization setups; it does not eliminate the need to validate an approximated policy before deployment.

Why it matters

PoEM approximated a target post-training policy from models already optimized on other rewards.

Limits and context

  • They further observe an approximately low-rank structure even when rewards are not linearly related, then estimate weights from reward or basis-policy outputs rather than launching another full training run.
  • PoEM predicts the outcome of the studied optimization setups; it does not eliminate the need to validate an approximated policy before deployment.

Key claims

  1. PoEM approximated a target post-training policy from models already optimized on other rewards.

    Qualification: They further observe an approximately low-rank structure even when rewards are not linearly related, then estimate weights from reward or basis-policy outputs rather than launching another full training run.

    Evidence: source-2026-09-26-005

Sources

  1. arXiv preprint 2609.30226arXiv · primary research

Corrections

No corrections have been recorded for this story.