frontier models
The Best Answer Was Hiding the Rest of the Search
Evolution-strategy post-training preserved broader pass-at-k solution coverage than reinforcement learning in the authors' discovery tests.

Summary
Evolution-strategy post-training preserved broader pass-at-k solution coverage than reinforcement learning in the authors' discovery tests.
The study asks whether optimizing a model around its highest-reward answers can narrow the range of solutions it will explore. Across the reported math and science settings, population-based evolution strategies produced a broader output distribution and consistently higher pass-at-k than reinforcement learning. The result supports diversity-oriented post-training where many valid candidates matter; it does not establish that evolution strategies dominate reinforcement learning for every objective.
Why it matters
Evolution-strategy post-training preserved broader pass-at-k solution coverage than reinforcement learning in the authors' discovery tests.
Limits and context
- The result supports diversity-oriented post-training where many valid candidates matter; it does not establish that evolution strategies dominate reinforcement learning for every objective.
Key claims
Evolution-strategy post-training preserved broader pass-at-k solution coverage than reinforcement learning in the authors' discovery tests.
Qualification: The result supports diversity-oriented post-training where many valid candidates matter; it does not establish that evolution strategies dominate reinforcement learning for every objective.
Evidence: source-2026-08-15-003
Sources
- arXiv preprint 2608.12679arXiv · primary research
Corrections
No corrections have been recorded for this story.