Skip to slide
Chapter 3 · Chatbot Arena, Letting Real People Pick the Winner
21 / 39

CHAPTER 03 · Chatbot Arena, Letting Real People Pick the Winner · 1 / 6

The problem it solves

A static benchmark with ground truth answers has three weaknesses when you care about open-ended chat quality. Its questions are fixed, so they cannot capture the messy, interactive way people really use a chatbot. Its fixed test set can leak into training data and become contaminated, quietly inflating scores. And for many real requests, like "help me word this difficult email," there simply is no single correct answer to grade against.

What you actually want to know is which model people prefer when they use it for real. That calls for a different kind of benchmark: one whose questions are live, always fresh from real users, and whose metric is human preference rather than a correct answer.

← → arrow keys work too