Skip to slide
Chapter 2 · SWE-bench, Can a Model Fix a Real Bug in Real Code?
19 / 39

CHAPTER 02 · SWE-bench, Can a Model Fix a Real Bug in Real Code? · 5 / 5

The one-sentence takeaway

SWE-bench measures real software ability by handing a model an actual GitHub bug and a whole codebase and then simply running the tests, turning "can it code?" into an honest, pass-or-fail question that top models initially failed more than 98 percent of the time.

Next: Chapter 3, Chatbot Arena, where there is no correct answer to check against, only which reply real people prefer.

← → arrow keys work too