CHAPTER 01 · BIG-Bench, Testing the Full Breadth of What a Model Can Do · 1 / 6
The problem it solves
By 2022, language models were improving so fast that the tests used to measure them kept becoming useless. A benchmark would come out, models would reach human level on it within a year or two, and then the score stopped being informative because everyone was near the top. This is called saturation: when a test is too easy, every model gets a high score and you can no longer tell them apart.
There was a second, subtler problem. Most benchmarks measured a narrow slice of ability, often one task like translation or sentiment. But a large model is a generalist, and a handful of narrow tests cannot tell you where its broad knowledge holds up and where it quietly falls apart. To prepare for models that were getting more capable in unpredictable ways, researchers wanted a task suite that was wide rather than deep: many different kinds of problem, including strange ones nobody had thought to test before.