Skip to slide
Chapter 3 · Alignment, Turning a Text Predictor Into a Helpful Assistant
19 / 74

CHAPTER 03 · Alignment, Turning a Text Predictor Into a Helpful Assistant

Alignment, Turning a Text Predictor Into a Helpful Assistant

Papers: Training Language Models to Follow Instructions with Human Feedback (InstructGPT / RLHF, 2022) and Direct Preference Optimization (DPO, 2023)

After Chapters 1 and 2, we have a big model trained on the internet. But here is a surprise: a freshly trained model is not a helpful assistant. It is a very good autocomplete. Its only skill is predicting the next word in internet text. Ask it "How do I bake bread?" and it might continue with more questions, because on the internet, questions are often followed by more questions. It does not know that you want an answer.

These two papers are about closing that gap. The goal is alignment: making the model do what humans actually want, helpfully and safely. This is the step that turned GPT-3 into ChatGPT.

← → arrow keys work too