Skip to slide
Chapter 4 · DeepSeek-R1, Learning to Reason Through Reinforcement Learning
32 / 53

CHAPTER 04 · DeepSeek-R1, Learning to Reason Through Reinforcement Learning · 5 / 6

Why this paper mattered

DeepSeek-R1 was a landmark for two reasons. First, it showed that high-level reasoning can be grown through reinforcement learning with simple verifiable rewards, rather than painstakingly taught example by example. That is a more scalable path, because checking answers is far cheaper than hand-writing reasoning. Second, by being open, it pulled back the curtain on how frontier reasoning models are built, accelerating the entire field. It is the clearest demonstration yet of this folder's central theme: let the model spend more effort thinking, and reward it for thinking well, and powerful reasoning follows.

← → arrow keys work too