benchmarks evals
The Correct Answer Survived Behind the Wrong Score
A two-parameter, label-free correction recovered decisions from hidden states after native sequence scoring collapsed.
Summary
A two-parameter, label-free correction recovered decisions from hidden states after native sequence scoring collapsed.
The study distinguishes absent reasoning from an output bottleneck: hidden-state probes decoded correct answers even when native sequence scores failed under structural biases. A minimal additive correction fitted on as few as 25 unlabeled examples recovered 9 to 34 accuracy points for Qwen3.5 models and transferred to OLMo-2-1B and Llama-3.1-8B. Hard-instance and permutation controls supported an instance-specific signal, but the result narrows how some benchmark failures should be interpreted rather than proving latent correctness in every error.
Why it matters
A two-parameter, label-free correction recovered decisions from hidden states after native sequence scoring collapsed.
Limits and context
No additional limitation was separately recorded.
Key claims
A two-parameter, label-free correction recovered decisions from hidden states after native sequence scoring collapsed.
Evidence: source-2026-09-01-005
Sources
- arXiv preprint 2608.31068arXiv · primary research
Corrections
No corrections have been recorded for this story.