Skip to slide
Chapter 4 · Glossary: Benchmarks
35 / 39

CHAPTER 04 · Glossary: Benchmarks · 8 / 12

Contamination

Contamination is when the questions or answers from a benchmark accidentally end up in a model's training data. The model can then score well by having effectively seen the test in advance, which is memorization, not skill. Because most benchmarks are published openly on the internet, contamination is a constant threat to honest evaluation, and it gets worse as training datasets scrape ever more of the web. Defenses include keeping a hidden test set, refreshing questions over time, and the canary-string trick described next.

← → arrow keys work too