CHAPTER 04 · Glossary: Benchmarks · 10 / 12
Scoring and scaling behavior
Calibration
Calibration measures whether a model's confidence matches its accuracy. A well-calibrated model that says it is 70 percent sure is right about 70 percent of the time. A badly calibrated model might be wrong half the time while sounding completely certain. Calibration matters because a confident wrong answer is far more dangerous than a hesitant one. BIG-Bench found that calibration improves as models grow larger, but was still poor in absolute terms for the models of its day.
Aggregate score
An aggregate score is a single number that summarizes performance across many tasks, usually by averaging. It is convenient, because it lets you rank models at a glance, but it can hide a lot: a strong average can mask total failure on a few important tasks, and the way you combine tasks affects the result. This is why broad benchmarks report not just one aggregate but also breakdowns by task and comparisons to a human baseline.
Breakthrough behavior
Breakthrough behavior describes a skill that barely improves as a model grows, then jumps sharply once the model crosses some critical size. Plotted against scale, the curve stays flat and then suddenly climbs, rather than rising smoothly. BIG-Bench found that tasks with breakthrough behavior tend to require several reasoning steps chained together, while tasks that improve smoothly tend to lean on knowledge or memorization. This is the same phenomenon as emergent ability, observed across hundreds of tasks at once. A practical warning attached to it: whether a task looks like a breakthrough can depend on the exact metric used, so these curves should be read with care.
Brittleness
Brittleness is the tendency of a model's score to swing on small, meaningless changes, such as rewording a question, reordering options, or changing formatting. A brittle model has not really mastered a task; it has latched onto surface cues. BIG-Bench documented that even very large models are brittle in this way, which is a key reason no single benchmark number should be trusted on its own.
BIG-Bench Lite
BIG-Bench Lite is a smaller, curated subset of the full BIG-Bench suite, chosen to be representative so that researchers can get a quick read on a model without the expense of running all 204 tasks. It reflects a common pattern in evaluation: pair a large, thorough benchmark with a cheaper proxy version for fast iteration, then run the full suite less often.