safety security
Clinicians Preferred Answers That Still Failed Safety Rubrics
More than 26,000 judgments showed that pairwise preference could hide specialty-specific clinical failure rates.
Summary
More than 26,000 judgments showed that pairwise preference could hide specialty-specific clinical failure rates.
Using 26,804 blinded pairwise judgments from more than 736 clinicians in over 28 countries, a preprint compared which model answer clinicians preferred with separate rubric scores for accuracy, harmlessness and other safety-critical qualities. Models that ranked well by preference still produced meaningful failures, and those failures varied across specialties. Surface features explained slightly more preference variation than differences in the safety rubrics. The authors propose reporting failure rates directly and adding clinically grounded adjustments rather than treating a single preference ranking as a safety measure.
Why it matters
More than 26,000 judgments showed that pairwise preference could hide specialty-specific clinical failure rates.
Limits and context
No additional limitation was separately recorded.
Key claims
More than 26,000 judgments showed that pairwise preference could hide specialty-specific clinical failure rates.
Evidence: source-2026-08-05-004
Sources
- arXiv preprint 2608.02617arXiv · primary research
Corrections
No corrections have been recorded for this story.