TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

A Search Engine Started Tracking the Tests Behind Model Scores

Benchmark Radar joins 1,283 source records with 12,916 numeric observations and preserves the evidence behind comparisons.

Published Updated Story ID: mp-2026-09-13-006
Read the complete editionStory JSON

Summary

Benchmark Radar joins 1,283 source records with 12,916 numeric observations and preserves the evidence behind comparisons.

Benchmark Radar is a living database for discovering AI evaluations, their datasets, code, model-card mentions and score histories. Its daily discovery system watches 37 first-party research and engineering sources. At release, the catalog contained 1,283 source records drawn from four benchmark catalogs and 12,916 numeric observations on 790 records. The project includes a web dashboard, saturation and trend views, downloadable evidence and an offline CLI. It is a new research infrastructure release, and the paper explicitly treats cross-setting score comparison as something to audit rather than assume.

Why it matters

Benchmark Radar joins 1,283 source records with 12,916 numeric observations and preserves the evidence behind comparisons.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. Benchmark Radar joins 1,283 source records with 12,916 numeric observations and preserves the evidence behind comparisons.

    Evidence: source-2026-09-13-006

Sources

  1. arXiv preprint 2609.11115arXiv · primary research

Corrections

No corrections have been recorded for this story.

A Search Engine Started Tracking the Tests Behind Model Scores · The Machine Press