CHAPTER 07 · Glossary: Foundational Modelling · 20 / 27
RLHF
RLHF stands for Reinforcement Learning from Human Feedback. It is the three-step recipe (Chapter 3) that turns a raw pre-trained model into a helpful assistant: first supervised fine-tuning on human-written answers, then training a reward model from human rankings, then using reinforcement learning (specifically PPO) to optimize the model against that reward.
RLHF was the breakthrough that made ChatGPT possible. Its central trick is learning human taste from comparisons rather than trying to write down rules for good behavior.