Skip to slide
Chapter 2 · SWE-bench, Can a Model Fix a Real Bug in Real Code?
17 / 39

CHAPTER 02 · SWE-bench, Can a Model Fix a Real Bug in Real Code? · 3 / 5

What they found

The first results were humbling, and that was the point. The best model they tested, Claude 2, resolved only about 1.96 percent of the issues, fewer than one in fifty. State-of-the-art models that looked dazzling on older coding tests could barely make a dent in real software work. That huge gap is exactly what a good frontier benchmark is supposed to expose: it leaves enormous room to improve and gives the field a clear, hard target.

Two practical contributions came alongside the benchmark. The authors released a training set, SWE-bench-train, of around 19,000 non-test task instances from 37 repositories, so that others could train models for this skill. Using it, they fine-tuned two open models, SWE-Llama 7b and 13b, which had to process long contexts of over 100,000 tokens to read enough of a codebase, and in some settings these were competitive with the much larger proprietary models.

← → arrow keys work too