frontier models
No Agent Model Won Every Kind of Work
DAREBench placed 233 tasks into six workload groups and compared 35 commercial and open models over 7,587 runs.
Summary
DAREBench placed 233 tasks into six workload groups and compared 35 commercial and open models over 7,587 runs.
DAREBench adapts tasks from 22 source benchmarks into a two-by-three matrix defined by input modality and execution form. All tasks run in a shared environment with contract-based scoring and evidence audits. Across 23 commercial API models and 12 locally deployed open-weight models, no system dominated all groups. Text and multimodal workloads produced different accuracy-cost trade-offs, while local models were competitive in several groups but trailed frontier commercial systems overall. The benchmark argues that deployment choices should follow workload profiles rather than a single aggregate score; its conclusions remain tied to the selected tasks, environment and reference cost assumptions.
Why it matters
DAREBench placed 233 tasks into six workload groups and compared 35 commercial and open models over 7,587 runs.
Limits and context
- The benchmark argues that deployment choices should follow workload profiles rather than a single aggregate score; its conclusions remain tied to the selected tasks, environment and reference cost assumptions.
Key claims
DAREBench placed 233 tasks into six workload groups and compared 35 commercial and open models over 7,587 runs.
Qualification: The benchmark argues that deployment choices should follow workload profiles rather than a single aggregate score; its conclusions remain tied to the selected tasks, environment and reference cost assumptions.
Evidence: source-2026-09-09-004
Sources
- arXiv preprint 2609.06059arXiv · primary research
Corrections
No corrections have been recorded for this story.