benchmarks evals
The Best Coding Agent Solved Fewer Than Half
SWE-bench Science spans 119 tasks, 98 repositories and twenty fields where software is part of the instrument.
Summary
SWE-bench Science spans 119 tasks, 98 repositories and twenty fields where software is part of the instrument.
The benchmark separates issue-driven, expert-exploratory and engineering-integration work and reports a best pass-at-one below 50 percent. Its error analysis finds failures in scientific abstraction, exploration, repair coverage and generalization; a paired ablation also showed that well-grounded scientific guidance can help while poorly aligned guidance can anchor the repair in the wrong direction.
Why it matters
SWE-bench Science spans 119 tasks, 98 repositories and twenty fields where software is part of the instrument.
Limits and context
No additional limitation was separately recorded.
Key claims
SWE-bench Science spans 119 tasks, 98 repositories and twenty fields where software is part of the instrument.
Evidence: source-2026-08-23-009
Sources
- arXiv preprint 2608.19799arXiv · primary research
Corrections
No corrections have been recorded for this story.