Skip to slide
Chapter 5 · Mixtral and Mixture of Experts, More Brain, Same Speed
33 / 74

CHAPTER 05 · Mixtral and Mixture of Experts, More Brain, Same Speed · 1 / 6

The tension we are trying to escape

Recall the trade-off from earlier chapters:

  • A bigger model knows more and reasons better.
  • But every time you use a normal model, all of its parameters do work, so a bigger model costs more for every single answer. This running cost is called inference.

A normal model is dense: the whole network fires for every word. MoE breaks this rule. It builds a model with a huge total number of parameters, but arranges things so that only a small slice of them activates for any given word. Lots of knowledge stored, little work done per word.

← → arrow keys work too