Skip to slide
Chapter 2 · SWE-bench, Can a Model Fix a Real Bug in Real Code?
14 / 39

CHAPTER 02 · SWE-bench, Can a Model Fix a Real Bug in Real Code?

SWE-bench, Can a Model Fix a Real Bug in Real Code?

Paper: SWE-bench: Can Language Models Resolve Real-World GitHub Issues? (2023)

BIG-Bench in the last chapter measured breadth with made-up tasks. SWE-bench takes the opposite tack: one domain, software engineering, but using completely real work pulled straight from open-source projects. It asks a model to do something a professional developer does every day, fix a reported bug in a large existing codebase, and it grades the result the only way that truly counts: by running the tests.

← → arrow keys work too