Skip to slide
Chapter 6 · Judging Models, How Do We Measure Quality?
44 / 74

CHAPTER 06 · Judging Models, How Do We Measure Quality? · 5 / 7

Be careful: judges have biases

The paper is honest about the ways an LLM judge can be fooled, and knowing these is important if you ever rely on one:

  • Position bias: the judge may favor whichever answer it sees first, regardless of quality. The fix is to swap the order and average.
  • Verbosity bias: the judge tends to prefer longer answers, even when a shorter one is better. Length can masquerade as quality.
  • Self-enhancement bias: a judge may favor answers written in its own style, subtly rating its own family of models more highly.
flowchart TD
    J[LLM judge] --> B1[Position bias<br/>prefers the first answer]
    J --> B2[Verbosity bias<br/>prefers longer answers]
    J --> B3[Self-enhancement bias<br/>prefers its own style]

These biases do not make LLM judges useless. They make them tools to use carefully, with tricks like swapping answer order and watching out for padding.

← → arrow keys work too