Skip to slide
Chapter 1 · BIG-Bench, Testing the Full Breadth of What a Model Can Do
10 / 39

CHAPTER 01 · BIG-Bench, Testing the Full Breadth of What a Model Can Do · 3 / 6

What they found

The headline results are a tour of how scale changes model behavior.

Performance and calibration both improved as models got bigger, but stayed poor in absolute terms and well below the human raters. In other words, scale helped, but these models were still far from human-level breadth in 2022.

Different model families behaved surprisingly similarly at the same size, though sparse models (the mixture-of-experts kind) got a bit more out of each unit of compute.

The most interesting finding was about how skills appear with scale. Some tasks improved smoothly and predictably as models grew, and these usually leaned on knowledge or memorization. Other tasks showed breakthrough behavior: almost no improvement for a long time, then a sudden jump once the model crossed a critical size. These breakthrough tasks tended to require several steps chained together. This is the same emergent ability idea you may have met before, measured here across hundreds of tasks at once.

They also found that large models are brittle: small, meaningless changes in how a task is worded could swing the score a lot, a warning that a single number never tells the whole story. And social bias often grew with scale when the prompt was ambiguous, though careful prompting could reduce it.

← → arrow keys work too