benchmarks evals
The Answer Key Recomputed Itself From Live Data
Executable ground truth improved expert-judge agreement 29% by MCC while reducing tokens 16% per test case.
Summary
Executable ground truth improved expert-judge agreement 29% by MCC while reducing tokens 16% per test case.
Static answers go stale when a data-science agent is asked about live audience or business data. This framework encodes the expected result as a function that recomputes directly from the underlying database at evaluation time, then scores atomic facts regardless of whether the agent replies in prose, a table or HTML. In an author-reported human agreement study using a synthetic database shaped like a production system, ground-truth-as-code improved Matthews correlation with expert labels 29% over natural-language references and used 16% fewer tokens. A self-directed baseline without explicit ground truth was anti-correlated with human judgment.
Why it matters
Executable ground truth improved expert-judge agreement 29% by MCC while reducing tokens 16% per test case.
Limits and context
No additional limitation was separately recorded.
Key claims
Executable ground truth improved expert-judge agreement 29% by MCC while reducing tokens 16% per test case.
Evidence: source-2026-09-16-009
Sources
- arXiv preprint 2609.16487arXiv · primary research
Corrections
No corrections have been recorded for this story.