safety
The Intake Bot Asked More but Inferred Too Much
A six-clinician pilot found more vignette facts and more unsupported clinical inferences from an AI interviewer.
Summary
A six-clinician pilot found more vignette facts and more unsupported clinical inferences from an AI interviewer.
A clinician-grounded test platform used simulated patient vignettes to compare open-ended psychiatric intake approaches. In a 25-minute pilot with six clinicians, the LLM interviewer recovered 88.0% of embedded clinically relevant items versus 38.9% for clinicians, but made unsupported clinical inferences more often, 56.8% versus 27.8%, and identified safety concerns less often, 33.3% versus 66.7%. The small simulated study is a quality-assurance probe, not evidence that the system is safe for patient use.
Why it matters
A six-clinician pilot found more vignette facts and more unsupported clinical inferences from an AI interviewer.
Limits and context
- The small simulated study is a quality-assurance probe, not evidence that the system is safe for patient use.
Key claims
A six-clinician pilot found more vignette facts and more unsupported clinical inferences from an AI interviewer.
Qualification: The small simulated study is a quality-assurance probe, not evidence that the system is safe for patient use.
Evidence: source-2026-09-21-004
Sources
- arXiv preprint 2609.21149arXiv · primary research
Corrections
No corrections have been recorded for this story.