CHAPTER 03 · Chatbot Arena, Letting Real People Pick the Winner
Chatbot Arena, Letting Real People Pick the Winner
Paper: Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference (2024)
The first two benchmarks both score answers against something objective: a known answer in BIG-Bench, a passing test in SWE-bench. But some of the most important things about a chatbot have no objective answer at all. Which of two helpful, well-written replies is better? That depends on human taste and judgment. Chatbot Arena tackles this head-on by turning evaluation into a live, crowd-powered tournament where the only judges are real users.