CHAPTER 02 · Scaling Laws and Chinchilla, How Big Should a Model Be? · 2 / 5
Paper 1: Scaling Laws, the discovery that bigger is predictably better
Before this paper, people knew bigger models tended to be better, but it felt like alchemy. The Scaling Laws paper showed something striking: the improvement is smooth and predictable.
Specifically, as you increase model size, data, or compute, the model's error (measured as loss, where lower is better) drops along a clean curve called a power law. When you plot it on the right kind of graph, it is almost a straight line over many orders of magnitude.
flowchart LR
A[More parameters] --> L[Lower loss]
B[More training data] --> L
C[More compute] --> L
L --> P[The drop is smooth<br/>and predictable]
Why this mattered so much: predictability removes risk. If a small experiment shows the curve, you can extrapolate and forecast how a model 100 times larger will perform before you spend the money to build it. That forecast is what gave teams the conviction to scale up to GPT-3 and beyond.
The original paper's takeaway leaned toward one conclusion: if you have more compute, spend most of it on a bigger model. As we will see, that advice was slightly off, and the next paper fixed it.