TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

Models Could Answer the Hardware Question. Most Could Not Build the Model

Top systems cleared 90 percent on performance reasoning, yet nearly every configuration averaged below 15 percent when asked to generate analytical model code.

Published Updated Story ID: mp-2026-09-07-007
Read the complete editionStory JSON

Summary

Top systems cleared 90 percent on performance reasoning, yet nearly every configuration averaged below 15 percent when asked to generate analytical model code.

PerfReasoning tests whether language models can compare workload mappings, predict off-chip traffic and calculate buffer requirements. The strongest closed models exceeded 90 percent on reasoning questions and the best open-weight model reached 82.4 percent. Model construction was much harder: one reported configuration exceeded 80 percent, while all others averaged below 15 percent and varied across runs. The benchmark exposes a gap between plausible architectural answers and executable performance models; it does not measure every form of systems engineering.

Why it matters

Top systems cleared 90 percent on performance reasoning, yet nearly every configuration averaged below 15 percent when asked to generate analytical model code.

Limits and context

  • The benchmark exposes a gap between plausible architectural answers and executable performance models; it does not measure every form of systems engineering.

Key claims

  1. Top systems cleared 90 percent on performance reasoning, yet nearly every configuration averaged below 15 percent when asked to generate analytical model code.

    Qualification: The benchmark exposes a gap between plausible architectural answers and executable performance models; it does not measure every form of systems engineering.

    Evidence: source-2026-09-07-007

Sources

  1. arXiv preprint 2609.04476arXiv · primary research

Corrections

No corrections have been recorded for this story.

Models Could Answer the Hardware Question. Most Could Not Build the Model · The Machine Press