Skip to slide
Chapter 1 · BIG-Bench, Testing the Full Breadth of What a Model Can Do
07 / 39

CHAPTER 01 · BIG-Bench, Testing the Full Breadth of What a Model Can Do

BIG-Bench, Testing the Full Breadth of What a Model Can Do

Paper: Beyond the Imitation Game: Quantifying and Extrapolating the Capabilities of Language Models (2022), known as BIG-Bench

This paper is the result of an unusual experiment in scale, not of the model but of the test. More than 450 researchers across 132 institutions came together to build a single giant benchmark of 204 tasks, deliberately chosen to be hard for the language models of the day. The goal was simple to state and hard to do: build a test broad and difficult enough to actually map the edges of what these models can and cannot do.

← → arrow keys work too