CHAPTER 02 · SWE-bench, Can a Model Fix a Real Bug in Real Code? · 1 / 5
The problem it solves
Coding benchmarks before SWE-bench, like the popular HumanEval, mostly asked a model to write a small self-contained function from a short description, the kind of problem you could solve in a few lines. Real software work looks nothing like that. Fixing a real bug means navigating a sprawling repository, understanding how functions in different files affect each other, and making a small, precise change in the right place. None of that is captured by writing one tidy function in isolation.
The authors also faced the benchmark-builder's eternal dilemma, the same one BIG-Bench wrestled with. A good test must be hard enough to challenge top models, but its answers must still be easy to check automatically. Open-ended essays are hard to grade; tiny puzzles are easy to grade but too simple. Software has a beautiful property that resolves this: code can pose an arbitrarily hard problem, yet a solution can be checked objectively by running unit tests and seeing whether they pass.