CHAPTER 04 · DeepSeek-R1, Learning to Reason Through Reinforcement Learning · 4 / 6
Sharing the ability: distillation
There is one more valuable contribution. The team used distillation to transfer R1's reasoning skill into a family of smaller models (ranging from tiny to large). Distillation means training a smaller "student" model to imitate the outputs of a larger "teacher" model. The result was a set of compact models that reason surprisingly well, making strong reasoning accessible on far more modest hardware.