Skip to slide
Chapter 4 · Glossary: Benchmarks
39 / 39

CHAPTER 04 · Glossary: Benchmarks · 12 / 12

Ranking by human votes

Pairwise comparison

A pairwise comparison is a judgment between exactly two options: shown answer A and answer B, which is better? It is much easier and more reliable for people to compare two things than to score one thing on an absolute scale, because "is this a 7 or an 8 out of 10?" is hard while "is A better than B?" is natural. Chatbot Arena is built entirely on pairwise comparisons, one vote per head-to-head matchup.

Crowdsourcing

Crowdsourcing means collecting many small contributions from a large, open group of ordinary people rather than from a handful of experts. Chatbot Arena crowdsources both its questions (real users type whatever they want) and its judgments (those same users vote). The benefit is scale and diversity that no expert panel could match; the challenge is noise, since individual votes are inconsistent, which is why heavy statistics are needed to extract a reliable signal.

Leaderboard

A leaderboard is a public ranking of models from best to worst according to some benchmark. Leaderboards focus attention and drive competition, for better and worse: they make progress visible, but they also create pressure to optimize for the specific benchmark. Chatbot Arena's leaderboard became one of the most referenced in the field because its live, human-preference design is hard to game by memorization.

Elo rating

Elo is a rating system, originally from chess, that estimates each player's strength from the outcomes of head-to-head matches. The full mechanics are explained in the folder 01 glossary. The key intuition for this folder is that beating a strong opponent raises your rating more than beating a weak one, and losing to a weak opponent costs you more. Chatbot Arena uses Elo-style ratings as an intuitive way to rank models from pairwise votes.

Bradley-Terry model

The Bradley-Terry model is a statistical method, dating to 1952, for estimating each competitor's underlying strength from a record of pairwise wins and losses. It is closely related to Elo but is a cleaner, more principled way to fit all the data at once and to attach confidence to the result. Chatbot Arena uses it to convert hundreds of thousands of noisy, lopsided votes into a stable ranking, along with a confidence interval around each model's score.

Confidence interval

A confidence interval is a range that expresses how uncertain an estimate is. Instead of saying a model's rating is exactly 1200, you say it is 1200 give or take 15, meaning the true value is very likely within that band. Confidence intervals are essential for honest leaderboards: if two models' intervals overlap heavily, the data cannot really tell them apart, and claiming one is better would be overreaching. Chatbot Arena reports these intervals so readers can see which ranking gaps are real and which are noise.

Active sampling

Active sampling is choosing which comparisons to run based on what would be most informative, rather than comparing everything at random. In Chatbot Arena, it means preferentially matching up models whose relative strength is still uncertain, the way a tournament scheduler arranges the games that will best clarify the standings. This makes the rankings converge faster and wastes fewer votes, while keeping the statistics valid.

← → arrow keys work too