CHAPTER 03 · Let's Verify Step by Step, Rewarding Good Reasoning
Let's Verify Step by Step, Rewarding Good Reasoning
Paper: Let's Verify Step by Step (2023)
Chapters 1 and 2 got the model to reason and act. But a hard question lurks underneath: when a model produces a long chain of reasoning, how do we judge it? The obvious answer is to check whether the final answer is right. This paper shows that the obvious answer is not the best one. It turns out that grading each step of the reasoning, rather than just the final result, makes models dramatically better at hard problems. This is a subtle idea with big consequences.