CHAPTER 07 · Glossary: Foundational Modelling · 18 / 27
Supervised fine-tuning (SFT)
Supervised fine-tuning is the first step of alignment (Chapter 3). Humans write high-quality example answers to many prompts, and the model is trained to imitate them. "Supervised" means the model learns from labeled examples of the correct behavior.
SFT teaches the model the basic shape of being helpful: when asked a question, give an answer. It is the foundation that later steps (RLHF or DPO) build on.