CHAPTER 02 · SWE-bench, Can a Model Fix a Real Bug in Real Code?
SWE-bench, Can a Model Fix a Real Bug in Real Code?
Paper: SWE-bench: Can Language Models Resolve Real-World GitHub Issues? (2023)
BIG-Bench in the last chapter measured breadth with made-up tasks. SWE-bench takes the opposite tack: one domain, software engineering, but using completely real work pulled straight from open-source projects. It asks a model to do something a professional developer does every day, fix a reported bug in a large existing codebase, and it grades the result the only way that truly counts: by running the tests.