Skip to slide
Chapter 1 · BIG-Bench, Testing the Full Breadth of What a Model Can Do
13 / 39

CHAPTER 01 · BIG-Bench, Testing the Full Breadth of What a Model Can Do · 6 / 6

The one-sentence takeaway

BIG-Bench answered fast-improving models with a fast-broadening test, proving that the most honest way to measure a generalist is a huge, diverse, deliberately-hard suite of tasks measured against a human baseline.

Next: Chapter 2, SWE-bench, where the test stops using made-up tasks and starts using real bugs from real software.

← → arrow keys work too