Skip to slide
Chapter 6 · Judging Models, How Do We Measure Quality?
42 / 74

CHAPTER 06 · Judging Models, How Do We Measure Quality? · 3 / 7

Tool 2: Chatbot Arena, let the crowd vote

Chatbot Arena takes a completely different approach: a live, public website where real people compare models head to head.

flowchart TD
    U[A person types one question] --> Two[It goes to two<br/>anonymous models, A and B]
    Two --> Ans[Both answers shown<br/>side by side, names hidden]
    Ans --> Vote[The person votes<br/>for the better answer]
    Vote --> Elo[Update each model's<br/>Elo rating]
    Elo --> Board[Public leaderboard]

Because the model names are hidden, the votes are unbiased by reputation. The votes feed an Elo rating, the same system used to rank chess players, where beating a strong opponent raises your score more than beating a weak one. With enough votes, this produces a remarkably trustworthy ranking grounded directly in human preference. The catch is that it is slow and expensive: it needs a constant stream of thousands of human voters.

← → arrow keys work too