TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

safety security

Clinicians Preferred Answers That Still Failed Safety Rubrics

More than 26,000 judgments showed that pairwise preference could hide specialty-specific clinical failure rates.

Published Updated Story ID: mp-2026-08-05-004
Read the complete editionStory JSON

Summary

More than 26,000 judgments showed that pairwise preference could hide specialty-specific clinical failure rates.

Using 26,804 blinded pairwise judgments from more than 736 clinicians in over 28 countries, a preprint compared which model answer clinicians preferred with separate rubric scores for accuracy, harmlessness and other safety-critical qualities. Models that ranked well by preference still produced meaningful failures, and those failures varied across specialties. Surface features explained slightly more preference variation than differences in the safety rubrics. The authors propose reporting failure rates directly and adding clinically grounded adjustments rather than treating a single preference ranking as a safety measure.

Why it matters

More than 26,000 judgments showed that pairwise preference could hide specialty-specific clinical failure rates.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. More than 26,000 judgments showed that pairwise preference could hide specialty-specific clinical failure rates.

    Evidence: source-2026-08-05-004

Sources

  1. arXiv preprint 2608.02617arXiv · primary research

Corrections

No corrections have been recorded for this story.