TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

Saying Confident Did Not Mean Being Calibrated

Across 30 models, verbal confidence frequently diverged from logits or semantic-entropy uncertainty.

Published Updated Story ID: mp-2026-08-31-012
Read the complete editionStory JSON

Summary

Across 30 models, verbal confidence frequently diverged from logits or semantic-entropy uncertainty.

The study compares linguistic confidence with internal signals across eight classification tasks and two generation tasks. Association was weak on average, instruction tuning often raised reported confidence while worsening calibration, and attitude cues inflated scores without improving alignment. Score exemplars sometimes preserved rank ordering, but the authors conclude that verbal confidence needs multi-axis evaluation before it enters reliability pipelines.

Why it matters

Across 30 models, verbal confidence frequently diverged from logits or semantic-entropy uncertainty.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. Across 30 models, verbal confidence frequently diverged from logits or semantic-entropy uncertainty.

    Evidence: source-2026-08-31-012

Sources

  1. arXiv preprint 2608.28382arXiv · primary research

Corrections

No corrections have been recorded for this story.