research
A Million Medical Questions Came From Clinicians’ Shared Images
ThoughtMed-1M pairs de-identified images with verified expert commentary; its trained model reached 85.4 percent macro accuracy across 42 benchmarks.
Summary
ThoughtMed-1M pairs de-identified images with verified expert commentary; its trained model reached 85.4 percent macro accuracy across 42 benchmarks.
The authors built a pipeline around de-identified medical images and commentaries shared on clinician-oriented social media, using a language model plus clinician-in-the-loop verification to create more than one million long-form visual question-answer pairs. A foundation model trained on that set, FOLTMed, achieved 85.4 percent macro accuracy across 42 medical VQA benchmarks and outscored comparison systems by three to five points on reported factuality and similarity measures. The work presents a scalable research dataset and benchmark result, not clinical approval, and its provenance and de-identification controls remain central to responsible reuse.
Why it matters
ThoughtMed-1M pairs de-identified images with verified expert commentary; its trained model reached 85.4 percent macro accuracy across 42 benchmarks.
Limits and context
- The work presents a scalable research dataset and benchmark result, not clinical approval, and its provenance and de-identification controls remain central to responsible reuse.
Key claims
ThoughtMed-1M pairs de-identified images with verified expert commentary; its trained model reached 85.4 percent macro accuracy across 42 benchmarks.
Qualification: The work presents a scalable research dataset and benchmark result, not clinical approval, and its provenance and de-identification controls remain central to responsible reuse.
Evidence: source-2026-09-09-010
Sources
- arXiv preprint 2609.06914arXiv · primary research
Corrections
No corrections have been recorded for this story.