Skip to slide
Chapter 3 · Chatbot Arena, Letting Real People Pick the Winner
24 / 39

CHAPTER 03 · Chatbot Arena, Letting Real People Pick the Winner · 4 / 6

A short worked example

Suppose a new model joins the Arena and, in its first matches, beats a model that everyone already agrees is excellent. Under an Elo-style or Bradley-Terry rating, that win counts for a lot, because defeating a strong opponent is strong evidence of strength, and the newcomer's rating jumps. If it had instead beaten a weak model, the rating would barely move, since that result was expected. Over thousands of such matches the ratings settle into an order that reflects real relative quality, with the confidence interval shrinking as more votes come in. This is why a single lucky win cannot crown a model: the system demands consistent wins against tough competition.

← → arrow keys work too