benchmarks evals
Saying Confident Did Not Mean Being Calibrated
Across 30 models, verbal confidence frequently diverged from logits or semantic-entropy uncertainty.
Summary
Across 30 models, verbal confidence frequently diverged from logits or semantic-entropy uncertainty.
The study compares linguistic confidence with internal signals across eight classification tasks and two generation tasks. Association was weak on average, instruction tuning often raised reported confidence while worsening calibration, and attitude cues inflated scores without improving alignment. Score exemplars sometimes preserved rank ordering, but the authors conclude that verbal confidence needs multi-axis evaluation before it enters reliability pipelines.
Why it matters
Across 30 models, verbal confidence frequently diverged from logits or semantic-entropy uncertainty.
Limits and context
No additional limitation was separately recorded.
Key claims
Across 30 models, verbal confidence frequently diverged from logits or semantic-entropy uncertainty.
Evidence: source-2026-08-31-012
Sources
- arXiv preprint 2608.28382arXiv · primary research
Corrections
No corrections have been recorded for this story.