CHAPTER 02 · SWE-bench, Can a Model Fix a Real Bug in Real Code? · 5 / 5
The one-sentence takeaway
SWE-bench measures real software ability by handing a model an actual GitHub bug and a whole codebase and then simply running the tests, turning "can it code?" into an honest, pass-or-fail question that top models initially failed more than 98 percent of the time.
Next: Chapter 3, Chatbot Arena, where there is no correct answer to check against, only which reply real people prefer.