CHAPTER 04 · LoRA, Fine-Tuning Giant Models on a Budget · 4 / 6
Why this is such a big deal in practice
One base model, many tiny adapters
Since the original model stays frozen and untouched, you can train a separate small adapter for each task and keep them all. Want the model to switch from legal writing to casual chat? Swap the adapter, not the whole model.
flowchart LR
Base[One frozen base model] --> A[+ Legal adapter]
Base --> B[+ Medical adapter]
Base --> C[+ Customer-support adapter]
A --> O1[Legal assistant]
B --> O2[Medical assistant]
C --> O3[Support assistant]
This is wonderfully efficient. Instead of storing ten full models, you store one base model plus ten tiny adapters.
No slowdown when running
A natural worry: does adding an adapter make the model slower to use? No. After training, the adapter's adjustment can be merged back into the original weights, producing a model that runs at exactly the original speed. You get the customization for free at inference time.
A simple analogy
Think of the giant pre-trained model as a printed textbook that is expensive to reprint. Normal fine-tuning rewrites the entire book for each new course. LoRA instead clips a thin set of sticky notes onto the relevant pages. The book underneath never changes, the sticky notes are cheap to make, and you can keep different sets of notes for different courses and swap them in seconds.