Skip to slide
Chapter 6 · Glossary: Planning and Reasoning
52 / 53

CHAPTER 06 · Glossary: Planning and Reasoning · 11 / 12

Learning to reason

Reinforcement learning

Reinforcement learning, or RL, is training by trial and error guided by rewards rather than by copying examples. The model tries something, receives a reward signal indicating how good the result was, and adjusts to earn more reward over time.

RL is the engine behind Chapter 4. Rather than being shown how to reason, DeepSeek-R1 was rewarded for reaching correct answers, and it gradually discovered reasoning strategies on its own because those strategies earned more reward. (The basic mechanics also appear in the folder 01 glossary.)

Reward and reward signal

The reward (or reward signal) is the number that tells a reinforcement learning model how good an outcome was, the feedback it is trying to maximize. The art of RL is often in choosing a good reward. A reward that is easy to fake leads to a model that games the system; a reward that is honest and hard to fake, like checking whether a math answer is actually correct, leads to genuine improvement. That honesty is exactly what verifiable rewards provide.

Verifiable rewards

Verifiable rewards are rewards that can be checked automatically and with certainty, such as whether a math answer is correct or whether code passes its tests. They are the key ingredient in Chapter 4. Because they are cheap to compute and impossible to fake, they let a model be trained with reinforcement learning at huge scale without needing a separate, foolable learned judge. The model simply gets rewarded when its answer is genuinely right.

Supervised fine-tuning (SFT)

Supervised fine-tuning is training a model to imitate labeled example answers, the straightforward "here is the correct response, copy this style" form of training. It contrasts with reinforcement learning, which uses rewards instead of fixed correct answers. In Chapter 4, DeepSeek-R1-Zero deliberately skipped SFT to see if reasoning could emerge from pure RL, and the full DeepSeek-R1 added just a little clean example data to tidy up presentation. (See the folder 01 glossary for the role of SFT in alignment.)

Cold-start data

Cold-start data is a small, curated set of clean, well-formatted examples used to warm up a model before the main training (here, before reinforcement learning) in Chapter 4. Its purpose is to give the model good habits of presentation from the start, so it reasons not only correctly but also readably. It is the fix that turned the rough-around-the-edges R1-Zero into the polished DeepSeek-R1.

Distillation

Distillation is training a smaller "student" model to imitate the outputs of a larger, more capable "teacher" model, transferring much of the teacher's skill into a more compact and cheaper-to-run package. In Chapter 4, the team distilled DeepSeek-R1's reasoning ability into a family of smaller models, making strong reasoning available on modest hardware.


← → arrow keys work too