CHAPTER 06 · Judging Models, How Do We Measure Quality? · 1 / 7
Why measuring chat models is genuinely hard
For a long time, models were tested with benchmarks made of questions that have one clearly correct answer, often multiple choice. That works for "What is the capital of France?" But modern assistants are judged on open-ended tasks: "Write me a polite email declining this meeting," or "Explain recursion to a ten-year-old." These have no single right answer. A good response must be helpful, clear, well-toned, and genuinely useful, all qualities that resist a simple answer key.
So how do you score something that has thousands of valid answers, each better or worse in fuzzy human ways? This paper offers two complementary tools and one provocative idea.