research
The Population Kept More Ways to Be Right
Evolution-strategy post-training improved first-answer accuracy while preserving broader reasoning coverage than GRPO in the reported comparisons.

Summary
Evolution-strategy post-training improved first-answer accuracy while preserving broader reasoning coverage than GRPO in the reported comparisons.
The study separates success on the most likely answer from coverage across multiple attempts. Its authors report that evolution strategies increased Pass@1 while retaining higher Pass@K than GRPO, whose sampled reasoning diversity narrowed during training. A sequential GRPO-then-ES schedule combined the two tendencies. The paper also found that task gains came from a sparse subset of larger parameter updates despite broad movement across the model. These are author-reported experiments and theory on selected models and tasks, not evidence that evolution strategies dominate every reasoning workload.
Why it matters
Evolution-strategy post-training improved first-answer accuracy while preserving broader reasoning coverage than GRPO in the reported comparisons.
Limits and context
- These are author-reported experiments and theory on selected models and tasks, not evidence that evolution strategies dominate every reasoning workload.
Key claims
Evolution-strategy post-training improved first-answer accuracy while preserving broader reasoning coverage than GRPO in the reported comparisons.
Qualification: These are author-reported experiments and theory on selected models and tasks, not evidence that evolution strategies dominate every reasoning workload.
Evidence: source-2026-08-29-001
Sources
- arXiv preprint 2608.27351arXiv · primary research
Corrections
No corrections have been recorded for this story.