CHAPTER 01 · BIG-Bench, Testing the Full Breadth of What a Model Can Do · 2 / 6
The core idea
BIG-Bench is, at heart, an enormous and varied collection of tasks plus a fair way to run every model through all of them. The tasks were crowdsourced from hundreds of contributors, so they draw on linguistics, child development, math, common-sense reasoning, biology, physics, social bias, software, and much more. Many were designed specifically to be beyond what current models could do, which is where the name comes from.
flowchart TD
A[450+ contributors] --> B[Submit diverse hard tasks]
B --> C["BIG-Bench: 204 tasks"]
C --> D[Run every model<br/>across all tasks]
C --> E[Human expert raters<br/>do the same tasks]
D --> F[Compare models to<br/>each other and to humans]
E --> F
Three design choices are worth understanding because they show what makes a benchmark trustworthy.
First, a human baseline. A team of human expert raters worked through the tasks too, so that model scores could be compared against how well people do, not just against other models. A score of 60 means very different things depending on whether humans get 65 or 99.
Second, a smaller curated subset called BIG-Bench Lite. Running every model on all 204 tasks is expensive, so the authors picked a representative slice that gives a quick read without the full cost. This is a common pattern: a big benchmark and a cheap proxy version of it.
Third, and cleverest, a canary string. The whole benchmark is published on the internet, which creates a trap: future models might be trained on the benchmark itself and then "pass" it by memorization rather than skill, a problem called contamination. To fight this, every BIG-Bench document contains a unique identifier string. Model builders can search their training data for that string and remove the benchmark, and researchers can check whether a model has secretly seen it.