developer tools
The Better Prompt Was Forty-Seven Percent Shorter
ESPO clustered errors before proposing diverse fixes, then used bootstrap stability to keep prompt search from rewarding brittle gains.
Summary
ESPO clustered errors before proposing diverse fixes, then used bootstrap stability to keep prompt search from rewarding brittle gains.
Across seven public NLP benchmarks, ESPO averaged 74.67 percent accuracy against 70.91 percent for GEPA while producing prompts of 1,004 rather than 1,878 characters. The method separates diagnosis, four proposal strategies and a stability-based selection stage. An ablation found that adding diversity without bootstrap selection reduced performance by 1.20 points. Cross-model gains were reported on four additional students, but the large per-task variation means the average should not be treated as a universal prompt-optimization guarantee.
Why it matters
ESPO clustered errors before proposing diverse fixes, then used bootstrap stability to keep prompt search from rewarding brittle gains.
Limits and context
- Cross-model gains were reported on four additional students, but the large per-task variation means the average should not be treated as a universal prompt-optimization guarantee.
Key claims
ESPO clustered errors before proposing diverse fixes, then used bootstrap stability to keep prompt search from rewarding brittle gains.
Qualification: Cross-model gains were reported on four additional students, but the large per-task variation means the average should not be treated as a universal prompt-optimization guarantee.
Evidence: source-2026-09-06-005
Sources
- arXiv preprint 2609.04197arXiv · primary research
Corrections
No corrections have been recorded for this story.