benchmarks evals
The Culinary Judge Scored Every Possible Plate Before the Models Arrived
FlavourBench freezes executable scores for all 56 three-ingredient portfolios in each task.

Summary
FlavourBench freezes executable scores for all 56 three-ingredient portfolios in each task.
The benchmark evaluated 27 frontier endpoints on 534 identical tasks, totaling 14,418 scored model-task cells without differential missingness. Its two independently compiled panels correlated at 0.89, and 101 of 351 model pairs were statistically resolved; the authors release prompts, raw responses, score maps and an offline verifier.
Why it matters
FlavourBench freezes executable scores for all 56 three-ingredient portfolios in each task.
Limits and context
No additional limitation was separately recorded.
Key claims
FlavourBench freezes executable scores for all 56 three-ingredient portfolios in each task.
Evidence: source-2026-08-24-008
Sources
- arXiv preprint 2608.20574arXiv · primary research
Corrections
No corrections have been recorded for this story.