CHAPTER 03 · Alignment, Turning a Text Predictor Into a Helpful Assistant · 4 / 5
Putting the two papers together
| Aspect | InstructGPT / RLHF | DPO |
|---|---|---|
| Core idea | Learn human taste, then practice against it | Learn directly from preferred vs rejected pairs |
| Reward model | Separate, must be trained | None, the model itself plays that role |
| Reinforcement learning loop | Yes, using PPO | No |
| Complexity | High, can be unstable | Low, more stable |
| Both rely on | Human preference data | Human preference data |
Notice the bottom row. Both methods are powered by the same fuel: humans comparing answers. The difference is only in the machinery that turns those comparisons into a better model.