CHAPTER 01 · The Transformer, the Engine Inside Every Modern Model · 3 / 6
The rest of the machine
Attention is the star, but a Transformer block has a few more parts. Here is the full flow for one block.
flowchart TD
In[Word vectors come in] --> Att[Multi-head self-attention<br/>mix in context from other words]
Att --> Add1[Add the input back in<br/>and normalize]
Add1 --> FF[Feed-forward network<br/>think harder about each word on its own]
FF --> Add2[Add the input back in<br/>and normalize]
Add2 --> Out[Improved word vectors go out]
Two new pieces appear here:
- The feed-forward network is a small processing step applied to each word separately. If attention is "gather information from neighbors," the feed-forward step is "now think about what you gathered."
- The "add the input back in" arrows are called residual connections. They let the original information flow straight through, so deep stacks of blocks do not lose the thread. This is a practical trick that makes very deep models trainable.
A real Transformer stacks many of these blocks on top of each other. Each layer refines the meaning a little more. Early layers catch simple patterns like grammar; later layers catch abstract ones like intent.
One more thing: word order
Because attention looks at all words at once, it has no built-in sense of order. "Dog bites man" and "man bites dog" would look identical. The fix is positional encoding, a small signal added to each word's embedding that tells the model where the word sits in the sequence. Now order is preserved.