CHAPTER 02 · SWE-bench, Can a Model Fix a Real Bug in Real Code? · 4 / 5
Why it mattered
SWE-bench changed what "good at coding" means. It moved the goalposts from writing isolated snippets to resolving real issues in real projects, judged by real tests, and in doing so it created one of the defining challenges for the wave of agentic coding systems that followed. The very low starting scores gave the field a long, meaningful ramp to climb, and the percentage of SWE-bench issues a system can resolve quickly became one of the most-watched numbers in AI. The two software-agent papers in folder 03, SWE-agent and OpenHands, exist largely to push this number up.