CHAPTER 05 · Mixtral and Mixture of Experts, More Brain, Same Speed
Mixtral and Mixture of Experts, More Brain, Same Speed
Paper: Mixtral of Experts (2024)
We end folder 01 with a clever architecture trick. Chapter 2 taught us that bigger models are smarter but more expensive to run. What if you could have the knowledge of a big model while paying the running cost of a small one? That is exactly what a Mixture of Experts, or MoE, delivers, and Mixtral is the open model that made the idea famous.