CHAPTER 02 · SWE-bench, Can a Model Fix a Real Bug in Real Code? · 2 / 5
The core idea
SWE-bench builds its tasks from the natural history of real projects on GitHub. Whenever developers fix a bug or add a feature, they often file an issue describing the problem, then later submit a pull request (a bundle of code changes) that fixes it, and that pull request usually comes with tests that confirm the fix works. SWE-bench harvests these matched pairs.
flowchart TD
A[Real GitHub issue<br/>describing a bug] --> B[Give model the issue<br/>plus a snapshot of the codebase]
B --> C[Model writes a patch<br/>a set of code changes]
C --> D[Apply patch to the codebase]
D --> E["Run the project's real tests"]
E --> F{Do the tests pass?}
F -->|Yes| G[Issue resolved]
F -->|No| H[Not resolved]
The model is handed two things: the text of the issue, and a snapshot of the codebase as it was just before the fix. Its job is to produce a patch, the exact set of edits that resolves the issue. Then SWE-bench applies that patch and runs the repository's own test suite. If the tests that the original human fix was meant to satisfy now pass, the model resolved the issue. This is called execution-based evaluation: the score comes from actually running the code, not from comparing the model's text to a reference answer.
The full benchmark contains 2,294 such task instances drawn from 12 popular Python projects. Because new issues and pull requests appear constantly, the benchmark can be refreshed over time with minimal human effort, which helps it resist going stale.