Skip to slide
Chapter 3 · Let's Verify Step by Step, Rewarding Good Reasoning
21 / 53

CHAPTER 03 · Let's Verify Step by Step, Rewarding Good Reasoning · 1 / 6

The trap of grading only the final answer

Suppose a student solves a long math problem and gets the right final number. Did they understand it? Maybe. Or maybe they made two mistakes that happened to cancel out, or they guessed. If you only check the final answer, you cannot tell the difference, and you might reward sloppy or lucky reasoning that will fail next time.

This is the difference between two ways of giving feedback, and it is the central idea of the paper.

  • Outcome supervision: reward the model based only on whether the final answer is correct.
  • Process supervision: reward the model based on whether each individual step of the reasoning is correct.
flowchart TD
    Sol[A multi-step solution] --> O{How do we grade it?}
    O -->|Outcome| Final[Check only the<br/>final answer]
    O -->|Process| Steps[Check every<br/>reasoning step]
    Final --> Risk[A lucky right answer<br/>from flawed steps<br/>still gets full marks]
    Steps --> Precise[Catches exactly<br/>where reasoning breaks]
← → arrow keys work too