CHAPTER 01 · BIG-Bench, Testing the Full Breadth of What a Model Can Do · 5 / 6
Why it mattered
BIG-Bench set the template for the modern broad evaluation. It showed that the right response to fast-improving models is not one clever test but a wide, collaborative, deliberately-hard suite, paired with a human baseline and defenses against contamination. Its findings about smooth versus breakthrough scaling shaped how the field thinks about emergent abilities, and its canary-string trick is now standard practice. Later benchmarks, including the two in this folder, inherited its core lesson: a benchmark is only as good as it is hard to fake and broad enough to be honest.