benchmarks evals
The Same Model Name Was Not the Same Instrument Tomorrow
In 52,988 audited requests, shared-endpoint judges missed preregistered repeatability thresholds within one window and across days.

Summary
In 52,988 audited requests, shared-endpoint judges missed preregistered repeatability thresholds within one window and across days.
Two preregistered campaigns first tested whether a language-model judge was stable enough to measure anything else. Same-window repeat rankings reached Spearman 0.400 against a required 0.90, while byte-identical next-day replays reached 0.78 against 0.99. The authors trace the failures to label mapping, differences far below the instrument's noise floor and changed rankings for identical inputs; waiting and switching among four providers did not repair the tested setup. The paper's conclusion is bounded to black-box observers on shared infrastructure: evaluate snapshot identity before freezing a gate.
Why it matters
In 52,988 audited requests, shared-endpoint judges missed preregistered repeatability thresholds within one window and across days.
Limits and context
- The authors trace the failures to label mapping, differences far below the instrument's noise floor and changed rankings for identical inputs; waiting and switching among four providers did not repair the tested setup.
Key claims
In 52,988 audited requests, shared-endpoint judges missed preregistered repeatability thresholds within one window and across days.
Qualification: The authors trace the failures to label mapping, differences far below the instrument's noise floor and changed rankings for identical inputs; waiting and switching among four providers did not repair the tested setup.
Evidence: source-2026-09-04-003
Sources
- arXiv preprint 2609.04198arXiv · primary research
Corrections
No corrections have been recorded for this story.