Skip to slide
Chapter 4 · DeepSeek-R1, Learning to Reason Through Reinforcement Learning
27 / 53

CHAPTER 04 · DeepSeek-R1, Learning to Reason Through Reinforcement Learning

DeepSeek-R1, Learning to Reason Through Reinforcement Learning

Paper: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (2025)

So far, models learned to reason mostly by imitation: we showed them step-by-step examples (Chapter 1) or human-labeled good and bad steps (Chapter 3). DeepSeek-R1 asks a bolder question. What if we do not show the model how to reason at all, and instead just reward it for getting answers right, letting it discover how to reason on its own? The result was an open-source model that rivals the best closed reasoning models, and a genuinely surprising finding about how reasoning can emerge.

← → arrow keys work too