TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

research

The Sparse Baseline Beat Retrieval on Clinical Concepts

TF-IDF reached 33.43% Recall@10; retrieval augmentation fell to 31.99% on masked SNOMED recommendations.

Published Updated Story ID: mp-2026-09-17-007
Read the complete editionStory JSON

Summary

TF-IDF reached 33.43% Recall@10; retrieval augmentation fell to 31.99% on masked SNOMED recommendations.

A benchmark built from 75,491 annotations in 272 MIMIC-IV discharge summaries masks a target mention and asks systems to rank SNOMED CT concepts from nearby clinical context. Sparse TF-IDF concept prototypes led the tested methods with 14.81% Recall@1 and 33.43% Recall@10. Retrieval augmentation did not improve it, reaching 31.99% Recall@10. Rare concepts remained the bottleneck: Recall@10 was 7.74% for concepts seen in one or two training notes, versus 43.90% when seen in more than ten; 9.66% of test pairs used concepts absent from training.

Why it matters

TF-IDF reached 33.43% Recall@10; retrieval augmentation fell to 31.99% on masked SNOMED recommendations.

Limits and context

  • Retrieval augmentation did not improve it, reaching 31.99% Recall@10.

Key claims

  1. TF-IDF reached 33.43% Recall@10; retrieval augmentation fell to 31.99% on masked SNOMED recommendations.

    Qualification: Retrieval augmentation did not improve it, reaching 31.99% Recall@10.

    Evidence: source-2026-09-17-007

Sources

  1. arXiv preprint 2609.17855arXiv · primary research

Corrections

No corrections have been recorded for this story.