TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

developer tools

The Verifier Changed What ‘Previous’ Meant

In a synthetic draft–verify–revise test, balanced accuracy ranged from 0.156 to near-perfect as models and reasoning settings changed.

Published Updated Story ID: mp-2026-09-14-006
Read the complete editionStory JSON

Summary

In a synthetic draft–verify–revise test, balanced accuracy ranged from 0.156 to near-perfect as models and reasoning settings changed.

Draft–verify–revise pipelines pass language from one model role to another, creating an opening for context-dependent words such as ‘previous’ to change referents. A study built ten base examples in three controlled conditions and tested six models across 21 reasoning configurations. Balanced accuracy ranged from 0.156 to near-perfect. GPT-5.2 rose from 0.156 without reasoning to 0.942 at its highest tested effort, while Gemini 3 Pro stayed above 0.94 across settings and achieved its low-effort result at roughly 5% of the reported per-trial cost of GPT-5.2 at xhigh. The narrow synthetic design supports an engineering warning, not a universal model ranking.

Why it matters

In a synthetic draft–verify–revise test, balanced accuracy ranged from 0.156 to near-perfect as models and reasoning settings changed.

Limits and context

  • The narrow synthetic design supports an engineering warning, not a universal model ranking.

Key claims

  1. In a synthetic draft–verify–revise test, balanced accuracy ranged from 0.156 to near-perfect as models and reasoning settings changed.

    Qualification: The narrow synthetic design supports an engineering warning, not a universal model ranking.

    Evidence: source-2026-09-14-006

Sources

  1. arXiv preprint 2609.12162arXiv · primary research

Corrections

No corrections have been recorded for this story.