CHAPTER 07 · Glossary: Foundational Modelling
Glossary: Foundational Modelling
This is your shared reference for folder 01. Every term that the chapters link to is explained here from scratch, in plain language. Unlike a normal glossary that gives one-line definitions, this one takes the time to make each idea actually click. You can read it straight through as a primer, or jump in whenever a chapter sends you here.
Terms are grouped by theme so related ideas sit together.
- The basic building blocks: Neural network, Parameters (weights), Layers and depth, Vector, Token, Tokenization, Embedding
- Inside the Transformer: Self-attention, Multi-head attention, Feed-forward network, Positional encoding, Softmax, Encoder and decoder, Transformer, Context window
- Training the model: Pre-training, Fine-tuning, Loss function, Gradient descent and backpropagation, Overfitting, Hyperparameter, Perplexity
- Scale and cost: Compute and FLOPs, Scaling law, Emergent ability, In-context learning, Inference
- Alignment: Supervised fine-tuning (SFT), Reinforcement learning, RLHF, Reward model, PPO, KL divergence, Preference data, DPO
- Efficiency tricks: Matrix rank (low-rank), LoRA, Mixture of Experts (MoE), Router (gating network), Sparse and dense models
- Measuring quality: Benchmark, Elo rating