TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

developer tools

The Better Prompt Was Forty-Seven Percent Shorter

ESPO clustered errors before proposing diverse fixes, then used bootstrap stability to keep prompt search from rewarding brittle gains.

Published Updated Story ID: mp-2026-09-06-005
Read the complete editionStory JSON

Summary

ESPO clustered errors before proposing diverse fixes, then used bootstrap stability to keep prompt search from rewarding brittle gains.

Across seven public NLP benchmarks, ESPO averaged 74.67 percent accuracy against 70.91 percent for GEPA while producing prompts of 1,004 rather than 1,878 characters. The method separates diagnosis, four proposal strategies and a stability-based selection stage. An ablation found that adding diversity without bootstrap selection reduced performance by 1.20 points. Cross-model gains were reported on four additional students, but the large per-task variation means the average should not be treated as a universal prompt-optimization guarantee.

Why it matters

ESPO clustered errors before proposing diverse fixes, then used bootstrap stability to keep prompt search from rewarding brittle gains.

Limits and context

  • Cross-model gains were reported on four additional students, but the large per-task variation means the average should not be treated as a universal prompt-optimization guarantee.

Key claims

  1. ESPO clustered errors before proposing diverse fixes, then used bootstrap stability to keep prompt search from rewarding brittle gains.

    Qualification: Cross-model gains were reported on four additional students, but the large per-task variation means the average should not be treated as a universal prompt-optimization guarantee.

    Evidence: source-2026-09-06-005

Sources

  1. arXiv preprint 2609.04197arXiv · primary research

Corrections

No corrections have been recorded for this story.