CHAPTER 07 · Glossary: Foundational Modelling · 19 / 27
Reinforcement learning
Reinforcement learning, or RL, is a style of training where a model learns by trial and error guided by rewards, rather than by copying labeled examples. The model tries an action, receives a reward signal indicating how good the outcome was, and adjusts to earn more reward over time. It is how you train an agent to play a game: not by showing it the perfect moves, but by rewarding it when it wins.
In language models, RL is used to push a model toward responses that score well according to human preferences (see RLHF). It also powers the reasoning models in folder 02, such as DeepSeek-R1.