Skip to slide
Chapter 4 · Glossary: Benchmarks
29 / 39

CHAPTER 04 · Glossary: Benchmarks · 2 / 12

Static vs live benchmark

This is the single most useful distinction in the folder. A static benchmark is a fixed list of questions written down in advance, the same set every time, like a printed exam. A live benchmark draws a fresh stream of questions from the real world as it runs, so no two runs use exactly the same inputs. Static benchmarks are cheap, repeatable, and easy to share, which is why most benchmarks are static (BIG-Bench and the core of SWE-bench are static). Their weakness is that a fixed question set can leak into training data and can fail to capture how people really use a model. Live benchmarks, like Chatbot Arena, stay fresh and are much harder to memorize in advance, but they are harder to run and to make reproducible. Neither is simply better; they answer different questions.

← → arrow keys work too