developer tools
Two Hundred Twenty-One Green Patches Still Failed Review
SWE-Gate separates passing functional tests from satisfying constraints recovered from real pull-request reviews.
Summary
SWE-Gate separates passing functional tests from satisfying constraints recovered from real pull-request reviews.
The benchmark contains 303 repository-level repairs across 75 open-source Python projects, each with separate functional and review-constraint tests plus noncompliant and gold patches. Under one shared agent scaffold and four model backends, 644 generated repairs passed functional tests; 221 of those violated the supplied review constraints. The result shows how functional-only scoring can overstate repair completeness in the tested corpus, while the released package makes the distinction reproducible.
Why it matters
SWE-Gate separates passing functional tests from satisfying constraints recovered from real pull-request reviews.
Limits and context
- The result shows how functional-only scoring can overstate repair completeness in the tested corpus, while the released package makes the distinction reproducible.
Key claims
SWE-Gate separates passing functional tests from satisfying constraints recovered from real pull-request reviews.
Qualification: The result shows how functional-only scoring can overstate repair completeness in the tested corpus, while the released package makes the distinction reproducible.
Evidence: source-2026-09-04-007
Sources
- arXiv preprint 2609.04167arXiv · primary research
Corrections
No corrections have been recorded for this story.