developer tools
The Verifier Changed What ‘Previous’ Meant
In a synthetic draft–verify–revise test, balanced accuracy ranged from 0.156 to near-perfect as models and reasoning settings changed.
Summary
In a synthetic draft–verify–revise test, balanced accuracy ranged from 0.156 to near-perfect as models and reasoning settings changed.
Draft–verify–revise pipelines pass language from one model role to another, creating an opening for context-dependent words such as ‘previous’ to change referents. A study built ten base examples in three controlled conditions and tested six models across 21 reasoning configurations. Balanced accuracy ranged from 0.156 to near-perfect. GPT-5.2 rose from 0.156 without reasoning to 0.942 at its highest tested effort, while Gemini 3 Pro stayed above 0.94 across settings and achieved its low-effort result at roughly 5% of the reported per-trial cost of GPT-5.2 at xhigh. The narrow synthetic design supports an engineering warning, not a universal model ranking.
Why it matters
In a synthetic draft–verify–revise test, balanced accuracy ranged from 0.156 to near-perfect as models and reasoning settings changed.
Limits and context
- The narrow synthetic design supports an engineering warning, not a universal model ranking.
Key claims
In a synthetic draft–verify–revise test, balanced accuracy ranged from 0.156 to near-perfect as models and reasoning settings changed.
Qualification: The narrow synthetic design supports an engineering warning, not a universal model ranking.
Evidence: source-2026-09-14-006
Sources
- arXiv preprint 2609.12162arXiv · primary research
Corrections
No corrections have been recorded for this story.