benchmarks evals
Models Could Answer the Hardware Question. Most Could Not Build the Model
Top systems cleared 90 percent on performance reasoning, yet nearly every configuration averaged below 15 percent when asked to generate analytical model code.
Summary
Top systems cleared 90 percent on performance reasoning, yet nearly every configuration averaged below 15 percent when asked to generate analytical model code.
PerfReasoning tests whether language models can compare workload mappings, predict off-chip traffic and calculate buffer requirements. The strongest closed models exceeded 90 percent on reasoning questions and the best open-weight model reached 82.4 percent. Model construction was much harder: one reported configuration exceeded 80 percent, while all others averaged below 15 percent and varied across runs. The benchmark exposes a gap between plausible architectural answers and executable performance models; it does not measure every form of systems engineering.
Why it matters
Top systems cleared 90 percent on performance reasoning, yet nearly every configuration averaged below 15 percent when asked to generate analytical model code.
Limits and context
- The benchmark exposes a gap between plausible architectural answers and executable performance models; it does not measure every form of systems engineering.
Key claims
Top systems cleared 90 percent on performance reasoning, yet nearly every configuration averaged below 15 percent when asked to generate analytical model code.
Qualification: The benchmark exposes a gap between plausible architectural answers and executable performance models; it does not measure every form of systems engineering.
Evidence: source-2026-09-07-007
Sources
- arXiv preprint 2609.04476arXiv · primary research
Corrections
No corrections have been recorded for this story.