TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

research

A Million Medical Questions Came From Clinicians’ Shared Images

ThoughtMed-1M pairs de-identified images with verified expert commentary; its trained model reached 85.4 percent macro accuracy across 42 benchmarks.

Published Updated Story ID: mp-2026-09-09-010
Read the complete editionStory JSON

Summary

ThoughtMed-1M pairs de-identified images with verified expert commentary; its trained model reached 85.4 percent macro accuracy across 42 benchmarks.

The authors built a pipeline around de-identified medical images and commentaries shared on clinician-oriented social media, using a language model plus clinician-in-the-loop verification to create more than one million long-form visual question-answer pairs. A foundation model trained on that set, FOLTMed, achieved 85.4 percent macro accuracy across 42 medical VQA benchmarks and outscored comparison systems by three to five points on reported factuality and similarity measures. The work presents a scalable research dataset and benchmark result, not clinical approval, and its provenance and de-identification controls remain central to responsible reuse.

Why it matters

ThoughtMed-1M pairs de-identified images with verified expert commentary; its trained model reached 85.4 percent macro accuracy across 42 benchmarks.

Limits and context

  • The work presents a scalable research dataset and benchmark result, not clinical approval, and its provenance and de-identification controls remain central to responsible reuse.

Key claims

  1. ThoughtMed-1M pairs de-identified images with verified expert commentary; its trained model reached 85.4 percent macro accuracy across 42 benchmarks.

    Qualification: The work presents a scalable research dataset and benchmark result, not clinical approval, and its provenance and de-identification controls remain central to responsible reuse.

    Evidence: source-2026-09-09-010

Sources

  1. arXiv preprint 2609.06914arXiv · primary research

Corrections

No corrections have been recorded for this story.