Skip to slide
Chapter 7 · Glossary: Foundational Modelling
69 / 74

CHAPTER 07 · Glossary: Foundational Modelling · 22 / 27

PPO

PPO, short for Proximal Policy Optimization, is the specific reinforcement learning algorithm used in the RLHF recipe. You do not need its math, only its role: it is the method that adjusts the model to earn higher scores from the reward model, while taking care not to change the model too drastically in any one step. The "proximal" in the name refers to that caution, staying close to the previous version to keep training stable.

← → arrow keys work too