Skip to slide
Chapter 7 · Glossary: Foundational Modelling
72 / 74

CHAPTER 07 · Glossary: Foundational Modelling · 25 / 27

DPO

DPO stands for Direct Preference Optimization (Chapter 3). It achieves the same goal as RLHF, aligning a model to human preferences, but without a separate reward model or a reinforcement learning loop. Instead, it trains the model directly on preference data with a single, stable objective: make preferred answers more likely and rejected ones less likely, while staying close to the original model.

Its guiding insight is that "your language model is secretly a reward model," meaning the reward step can be folded mathematically into the model's own training. DPO is simpler, cheaper, and more stable than full RLHF, which is why it became so widely used.


← → arrow keys work too