CHAPTER 07 · Glossary: Foundational Modelling · 27 / 27
Measuring quality
Benchmark
A benchmark is a standardized test used to measure and compare model capabilities. Older benchmarks often used questions with a single correct answer, which made scoring easy but did not capture open-ended skills like helpfulness or writing quality.
Chapter 6 is about the harder problem of benchmarking open-ended chat ability, where there is no single right answer, and introduces tools like MT-Bench, Chatbot Arena, and using a strong model as an automatic judge.
Elo rating
Elo is a rating system originally designed to rank chess players, repurposed in Chapter 6 to rank chat models. Each model has a numeric rating. When two models compete and humans vote for the winner, the winner's rating goes up and the loser's goes down. Crucially, beating a strong opponent raises your rating more than beating a weak one, and losing to a weak opponent costs you more.
In Chatbot Arena, thousands of blind head-to-head human votes feed an Elo leaderboard, producing a ranking grounded directly in human preference. It is considered one of the most trustworthy ways to compare models, precisely because it is built on real human choices.