TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Hidden State Knew When the Vote Was Wrong

A leakage-free decodability score predicted when selecting an answer from internal signals would beat majority voting.

Published Updated Story ID: mp-2026-08-19-014
Read the complete editionStory JSON

Summary

A leakage-free decodability score predicted when selecting an answer from internal signals would beat majority voting.

CASE trains a linear gate on answer-token hidden states and chooses the highest-ranked sample. Its decodability measure predicted the gain over voting with a reported correlation of 0.75; across general and medical models, selection improved medium-difficulty accuracy by up to 19 points and hard questions by 16.8 points. A conventional probe looked strong only because question identity leaked across evaluation groups, underscoring that the criterion must itself be tested without leakage.

Why it matters

A leakage-free decodability score predicted when selecting an answer from internal signals would beat majority voting.

Limits and context

  • A conventional probe looked strong only because question identity leaked across evaluation groups, underscoring that the criterion must itself be tested without leakage.

Key claims

  1. A leakage-free decodability score predicted when selecting an answer from internal signals would beat majority voting.

    Qualification: A conventional probe looked strong only because question identity leaked across evaluation groups, underscoring that the criterion must itself be tested without leakage.

    Evidence: source-2026-08-19-014

Sources

  1. arXiv preprint 2608.17124arXiv · primary research

Corrections

No corrections have been recorded for this story.