TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Culinary Judge Scored Every Possible Plate Before the Models Arrived

FlavourBench freezes executable scores for all 56 three-ingredient portfolios in each task.

Published Updated Story ID: mp-2026-08-24-008
Read the complete editionStory JSON

Summary

FlavourBench freezes executable scores for all 56 three-ingredient portfolios in each task.

The benchmark evaluated 27 frontier endpoints on 534 identical tasks, totaling 14,418 scored model-task cells without differential missingness. Its two independently compiled panels correlated at 0.89, and 101 of 351 model pairs were statistically resolved; the authors release prompts, raw responses, score maps and an offline verifier.

Why it matters

FlavourBench freezes executable scores for all 56 three-ingredient portfolios in each task.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. FlavourBench freezes executable scores for all 56 three-ingredient portfolios in each task.

    Evidence: source-2026-08-24-008

Sources

  1. arXiv preprint 2608.20574arXiv · primary research

Corrections

No corrections have been recorded for this story.