CHAPTER 01 · BIG-Bench, Testing the Full Breadth of What a Model Can Do · 4 / 6
A short worked example
One BIG-Bench task is "checkmate-in-one": given a chess position, name the single move that delivers checkmate. It is a good illustration of why breadth matters. A model can be fluent and well-read yet fail this completely, because it requires following the rules of chess for several pieces at once rather than recalling a fact. Tasks like this are exactly the kind that show breakthrough behavior, staying near zero until a model is large enough to track the multiple steps involved. One narrow language test would never have revealed that gap; a suite of 204 varied tasks does.