TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Rubric Became a Program

ExecRubrics compiles evaluation logic into inspectable scoring functions instead of leaving every criterion to a black-box judge.

Published Updated Story ID: mp-2026-08-25-010
Read the complete editionStory JSON

Summary

ExecRubrics compiles evaluation logic into inspectable scoring functions instead of leaving every criterion to a black-box judge.

Across HealthBench, HelpSteer and ArgQuality, executable rubrics matched or improved natural-language rubric baselines at best preference accuracies of 53, 78 and 92 percent while cutting latency by as much as 320 times. The approach makes dependencies, penalties and override conditions explicit and editable.

Why it matters

ExecRubrics compiles evaluation logic into inspectable scoring functions instead of leaving every criterion to a black-box judge.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. ExecRubrics compiles evaluation logic into inspectable scoring functions instead of leaving every criterion to a black-box judge.

    Evidence: source-2026-08-25-010

Sources

  1. arXiv preprint 2608.22559arXiv · primary research

Corrections

No corrections have been recorded for this story.