frontier models
A New Reward Didn’t Need a New RL Run
PoEM approximated a target post-training policy from models already optimized on other rewards.
Summary
PoEM approximated a target post-training policy from models already optimized on other rewards.
The authors show that when a new reward is a linear combination of known rewards, its reinforcement-learned policy can also be combined in log-policy space. They further observe an approximately low-rank structure even when rewards are not linearly related, then estimate weights from reward or basis-policy outputs rather than launching another full training run. Experiments span synthetic and real rewards in text and image settings. PoEM predicts the outcome of the studied optimization setups; it does not eliminate the need to validate an approximated policy before deployment.
Why it matters
PoEM approximated a target post-training policy from models already optimized on other rewards.
Limits and context
- They further observe an approximately low-rank structure even when rewards are not linearly related, then estimate weights from reward or basis-policy outputs rather than launching another full training run.
- PoEM predicts the outcome of the studied optimization setups; it does not eliminate the need to validate an approximated policy before deployment.
Key claims
PoEM approximated a target post-training policy from models already optimized on other rewards.
Qualification: They further observe an approximately low-rank structure even when rewards are not linearly related, then estimate weights from reward or basis-policy outputs rather than launching another full training run.
Evidence: source-2026-09-26-005
Sources
- arXiv preprint 2609.30226arXiv · primary research
Corrections
No corrections have been recorded for this story.