Skip to slide
Chapter 4 · DeepSeek-R1, Learning to Reason Through Reinforcement Learning
28 / 53

CHAPTER 04 · DeepSeek-R1, Learning to Reason Through Reinforcement Learning · 1 / 6

The setup: reward correct answers, nothing else

The team started from a base model and applied reinforcement learning, the trial-and-error training you met in folder 01. But here is the twist that makes it work for reasoning. For problems like math and coding, you can check an answer automatically and for certain: a math answer is either correct or not, code either passes the tests or not. These are verifiable rewards, and they are powerful because they are cheap and impossible to fake.

This sidesteps the whole machinery of folder 01's alignment. There is no separate learned reward model that might be fooled. The reward comes straight from checking whether the answer is right.

flowchart TD
    M[Model attempts a problem] --> Ans[Produces reasoning<br/>and a final answer]
    Ans --> Check{Is the answer<br/>verifiably correct?}
    Check -->|Yes| Reward[Reward the model]
    Check -->|No| NoReward[No reward]
    Reward --> Better[Model adjusts to<br/>reason in ways that<br/>lead to correct answers]
    NoReward --> Better
    Better --> M
← → arrow keys work too