Skip to slide
Chapter 7 · Glossary: Foundational Modelling
55 / 74

CHAPTER 07 · Glossary: Foundational Modelling · 8 / 27

Self-attention

Self-attention is the core mechanism of the Transformer and arguably the most important idea in this whole folder. It lets every word in a sentence look at every other word and decide how much each one matters for understanding it.

Here is the intuition. Each word forms a question (a Query) about what it needs, and every word also advertises what it offers (a Key) and what it will contribute (a Value). A word's new meaning is built by blending in the Values of the words whose Keys best match its Query. So in "the trophy did not fit because it was too big," the word "it" can attend strongly to "trophy" and absorb that meaning.

The power of self-attention is that it connects distant words directly, in a single step, no matter how far apart they sit. This solved the long-range memory problem that older models struggled with. Chapter 1 walks through it in detail.

← → arrow keys work too