Skip to slide
Chapter 5 · Mixtral and Mixture of Experts, More Brain, Same Speed
32 / 74

CHAPTER 05 · Mixtral and Mixture of Experts, More Brain, Same Speed

Mixtral and Mixture of Experts, More Brain, Same Speed

Paper: Mixtral of Experts (2024)

We end folder 01 with a clever architecture trick. Chapter 2 taught us that bigger models are smarter but more expensive to run. What if you could have the knowledge of a big model while paying the running cost of a small one? That is exactly what a Mixture of Experts, or MoE, delivers, and Mixtral is the open model that made the idea famous.

← → arrow keys work too