TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

developer tools

Two Hundred Twenty-One Green Patches Still Failed Review

SWE-Gate separates passing functional tests from satisfying constraints recovered from real pull-request reviews.

Published Updated Story ID: mp-2026-09-04-007
Read the complete editionStory JSON

Summary

SWE-Gate separates passing functional tests from satisfying constraints recovered from real pull-request reviews.

The benchmark contains 303 repository-level repairs across 75 open-source Python projects, each with separate functional and review-constraint tests plus noncompliant and gold patches. Under one shared agent scaffold and four model backends, 644 generated repairs passed functional tests; 221 of those violated the supplied review constraints. The result shows how functional-only scoring can overstate repair completeness in the tested corpus, while the released package makes the distinction reproducible.

Why it matters

SWE-Gate separates passing functional tests from satisfying constraints recovered from real pull-request reviews.

Limits and context

  • The result shows how functional-only scoring can overstate repair completeness in the tested corpus, while the released package makes the distinction reproducible.

Key claims

  1. SWE-Gate separates passing functional tests from satisfying constraints recovered from real pull-request reviews.

    Qualification: The result shows how functional-only scoring can overstate repair completeness in the tested corpus, while the released package makes the distinction reproducible.

    Evidence: source-2026-09-04-007

Sources

  1. arXiv preprint 2609.04167arXiv · primary research

Corrections

No corrections have been recorded for this story.