TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

research

The Population Kept More Ways to Be Right

Evolution-strategy post-training improved first-answer accuracy while preserving broader reasoning coverage than GRPO in the reported comparisons.

Published Updated Story ID: mp-2026-08-29-001
Read the complete editionStory JSON

Summary

Evolution-strategy post-training improved first-answer accuracy while preserving broader reasoning coverage than GRPO in the reported comparisons.

The study separates success on the most likely answer from coverage across multiple attempts. Its authors report that evolution strategies increased Pass@1 while retaining higher Pass@K than GRPO, whose sampled reasoning diversity narrowed during training. A sequential GRPO-then-ES schedule combined the two tendencies. The paper also found that task gains came from a sparse subset of larger parameter updates despite broad movement across the model. These are author-reported experiments and theory on selected models and tasks, not evidence that evolution strategies dominate every reasoning workload.

Why it matters

Evolution-strategy post-training improved first-answer accuracy while preserving broader reasoning coverage than GRPO in the reported comparisons.

Limits and context

  • These are author-reported experiments and theory on selected models and tasks, not evidence that evolution strategies dominate every reasoning workload.

Key claims

  1. Evolution-strategy post-training improved first-answer accuracy while preserving broader reasoning coverage than GRPO in the reported comparisons.

    Qualification: These are author-reported experiments and theory on selected models and tasks, not evidence that evolution strategies dominate every reasoning workload.

    Evidence: source-2026-08-29-001

Sources

  1. arXiv preprint 2608.27351arXiv · primary research

Corrections

No corrections have been recorded for this story.