benchmarks evals
A Search Engine Started Tracking the Tests Behind Model Scores
Benchmark Radar joins 1,283 source records with 12,916 numeric observations and preserves the evidence behind comparisons.
Summary
Benchmark Radar joins 1,283 source records with 12,916 numeric observations and preserves the evidence behind comparisons.
Benchmark Radar is a living database for discovering AI evaluations, their datasets, code, model-card mentions and score histories. Its daily discovery system watches 37 first-party research and engineering sources. At release, the catalog contained 1,283 source records drawn from four benchmark catalogs and 12,916 numeric observations on 790 records. The project includes a web dashboard, saturation and trend views, downloadable evidence and an offline CLI. It is a new research infrastructure release, and the paper explicitly treats cross-setting score comparison as something to audit rather than assume.
Why it matters
Benchmark Radar joins 1,283 source records with 12,916 numeric observations and preserves the evidence behind comparisons.
Limits and context
No additional limitation was separately recorded.
Key claims
Benchmark Radar joins 1,283 source records with 12,916 numeric observations and preserves the evidence behind comparisons.
Evidence: source-2026-09-13-006
Sources
- arXiv preprint 2609.11115arXiv · primary research
Corrections
No corrections have been recorded for this story.