TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Answer Key Recomputed Itself From Live Data

Executable ground truth improved expert-judge agreement 29% by MCC while reducing tokens 16% per test case.

Published Updated Story ID: mp-2026-09-16-009
Read the complete editionStory JSON

Summary

Executable ground truth improved expert-judge agreement 29% by MCC while reducing tokens 16% per test case.

Static answers go stale when a data-science agent is asked about live audience or business data. This framework encodes the expected result as a function that recomputes directly from the underlying database at evaluation time, then scores atomic facts regardless of whether the agent replies in prose, a table or HTML. In an author-reported human agreement study using a synthetic database shaped like a production system, ground-truth-as-code improved Matthews correlation with expert labels 29% over natural-language references and used 16% fewer tokens. A self-directed baseline without explicit ground truth was anti-correlated with human judgment.

Why it matters

Executable ground truth improved expert-judge agreement 29% by MCC while reducing tokens 16% per test case.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. Executable ground truth improved expert-judge agreement 29% by MCC while reducing tokens 16% per test case.

    Evidence: source-2026-09-16-009

Sources

  1. arXiv preprint 2609.16487arXiv · primary research

Corrections

No corrections have been recorded for this story.