CHAPTER 02 · Scaling Laws and Chinchilla, How Big Should a Model Be?
Scaling Laws and Chinchilla, How Big Should a Model Be?
Papers: Scaling Laws for Neural Language Models (2020) and Training Compute-Optimal Large Language Models (Chinchilla, 2022)
Once you have the Transformer engine from Chapter 1, an obvious question appears: how big should you build it, and how much text should you feed it? These two papers answer that question with numbers instead of guesswork. Together they are the reason teams felt confident spending tens of millions of dollars to train ever-larger models.