Skip to slide
Chapter 3 · Let's Verify Step by Step, Rewarding Good Reasoning
22 / 53

CHAPTER 03 · Let's Verify Step by Step, Rewarding Good Reasoning · 2 / 6

Two kinds of graders

To put this into practice, the researchers trained two kinds of grading models, both relatives of the reward model you met in folder 01.

The PRM is a step-by-step verifier: a model whose job is to check reasoning, not to produce it. Human labelers went through thousands of solutions marking each step as correct or not, and the PRM learned to imitate those judgments.

← → arrow keys work too