research
The Sparse Baseline Beat Retrieval on Clinical Concepts
TF-IDF reached 33.43% Recall@10; retrieval augmentation fell to 31.99% on masked SNOMED recommendations.
Summary
TF-IDF reached 33.43% Recall@10; retrieval augmentation fell to 31.99% on masked SNOMED recommendations.
A benchmark built from 75,491 annotations in 272 MIMIC-IV discharge summaries masks a target mention and asks systems to rank SNOMED CT concepts from nearby clinical context. Sparse TF-IDF concept prototypes led the tested methods with 14.81% Recall@1 and 33.43% Recall@10. Retrieval augmentation did not improve it, reaching 31.99% Recall@10. Rare concepts remained the bottleneck: Recall@10 was 7.74% for concepts seen in one or two training notes, versus 43.90% when seen in more than ten; 9.66% of test pairs used concepts absent from training.
Why it matters
TF-IDF reached 33.43% Recall@10; retrieval augmentation fell to 31.99% on masked SNOMED recommendations.
Limits and context
- Retrieval augmentation did not improve it, reaching 31.99% Recall@10.
Key claims
TF-IDF reached 33.43% Recall@10; retrieval augmentation fell to 31.99% on masked SNOMED recommendations.
Qualification: Retrieval augmentation did not improve it, reaching 31.99% Recall@10.
Evidence: source-2026-09-17-007
Sources
- arXiv preprint 2609.17855arXiv · primary research
Corrections
No corrections have been recorded for this story.