Skip to slide
Chapter 2 · Scaling Laws and Chinchilla, How Big Should a Model Be?
16 / 74

CHAPTER 02 · Scaling Laws and Chinchilla, How Big Should a Model Be? · 3 / 5

Paper 2: Chinchilla, the correction

Two years later, a DeepMind team revisited the question with more careful experiments and found that the field had been building models that were too big and trained on too little data.

Here is the key finding in one line: for a given compute budget, you should grow the model and the dataset in equal measure. A useful rule of thumb that came out of this work is roughly 20 tokens of training data for every parameter.

To prove it, they trained a model called Chinchilla with 70 billion parameters on 1.4 trillion tokens. They compared it against Gopher, a model with 280 billion parameters (four times larger) trained on less data. Chinchilla, despite being four times smaller, beat the bigger model on almost everything.

flowchart TD
    Budget[Fixed compute budget] --> Q{How to spend it?}
    Q -->|Old way| Big[Huge model,<br/>not enough data<br/>example: 280B model]
    Q -->|Chinchilla way| Bal[Balanced model and data<br/>example: 70B model,<br/>about 20 tokens per parameter]
    Big --> Worse[Under-trained,<br/>weaker, expensive to run]
    Bal --> Better[Better results<br/>and cheaper to run]

A simple analogy

Imagine compute is a fixed budget for opening a restaurant. Model size is the size of the kitchen. Data is the amount of cooking practice your chefs get.

The old approach built a giant kitchen but barely let the chefs practice. Chinchilla showed that a medium kitchen with well-practiced chefs produces better food for the same total budget. A bigger kitchen is wasted if no one has learned to cook in it.

Why a smaller, well-trained model is a double win

A Chinchilla-style model is not only better, it is also cheaper to run. Every time you use a model (called inference), the cost depends on its size. A 70 billion parameter model costs far less per use than a 280 billion one. So training in a balanced way gives you a model that is both stronger and cheaper to serve to millions of users. That is why this paper reshaped how essentially every modern model is trained.

← → arrow keys work too