TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Best Coding Agent Solved Fewer Than Half

SWE-bench Science spans 119 tasks, 98 repositories and twenty fields where software is part of the instrument.

Published Updated Story ID: mp-2026-08-23-009
Read the complete editionStory JSON

Summary

SWE-bench Science spans 119 tasks, 98 repositories and twenty fields where software is part of the instrument.

The benchmark separates issue-driven, expert-exploratory and engineering-integration work and reports a best pass-at-one below 50 percent. Its error analysis finds failures in scientific abstraction, exploration, repair coverage and generalization; a paired ablation also showed that well-grounded scientific guidance can help while poorly aligned guidance can anchor the repair in the wrong direction.

Why it matters

SWE-bench Science spans 119 tasks, 98 repositories and twenty fields where software is part of the instrument.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. SWE-bench Science spans 119 tasks, 98 repositories and twenty fields where software is part of the instrument.

    Evidence: source-2026-08-23-009

Sources

  1. arXiv preprint 2608.19799arXiv · primary research

Corrections

No corrections have been recorded for this story.