TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

Physicians Mistook Synthetic Charts for Real Ones

The open benchmark contains 1,268 longitudinal patients and 5,602 fully synthetic encounters.

Published Updated Story ID: mp-2026-09-25-005
Read the complete editionStory JSON

Summary

The open benchmark contains 1,268 longitudinal patients and 5,602 fully synthetic encounters.

Synthetic Hospital builds longitudinal electronic-health-record cases from public medical-education material, with ontology-grounded diagnoses, findings and time relations rather than protected patient data. In a blinded review, physicians distinguished its charts from real records at a near-chance 53%. The best of ten tested models scored 0.73 severity-weighted F1 on reconstructing problem lists—matching the seven-physician mean on a subset but below the best physician's 0.89—and missed roughly half of clinically relevant findings in chart summaries.

Why it matters

The open benchmark contains 1,268 longitudinal patients and 5,602 fully synthetic encounters.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. The open benchmark contains 1,268 longitudinal patients and 5,602 fully synthetic encounters.

    Evidence: source-2026-09-25-005

Sources

  1. arXiv preprint 2609.30027arXiv · primary research

Corrections

No corrections have been recorded for this story.