Skip to slide
Chapter 2 · Scaling Laws and Chinchilla, How Big Should a Model Be?
14 / 74

CHAPTER 02 · Scaling Laws and Chinchilla, How Big Should a Model Be? · 1 / 5

Three knobs you can turn

When training a language model, you mostly control three things:

  1. Model size, the number of parameters. Parameters are the adjustable dials inside the model. More dials means more capacity to store patterns.
  2. Data size, the number of tokens of text you train on. A token is roughly a word or a piece of a word.
  3. Compute, the total amount of calculation you do, measured in FLOPs. Compute is basically time multiplied by hardware, which translates into money.

These three are linked. More compute lets you train a bigger model, or train on more data, or both. The whole game is deciding how to spend a fixed compute budget.

← → arrow keys work too