CHAPTER 04 · LoRA, Fine-Tuning Giant Models on a Budget · 3 / 6
How LoRA works
Instead of editing the original weights, LoRA does this:
- Freeze the original model completely. Not a single original parameter changes.
- Add a small pair of skinny matrices (the low-rank adapter) alongside the layers you want to adapt.
- Train only those small matrices on your new data. They learn the adjustment.
flowchart TD
Frozen[Original giant model<br/>FROZEN, never changes] --> Combine
Adapter[Tiny LoRA adapter<br/>the only thing that trains] --> Combine
Combine[Add the adapter's adjustment<br/>on top of the frozen model] --> Output[Model specialized<br/>for your task]
Because you are training only the skinny matrices, the number of trainable parameters can drop by a factor of thousands. Memory needs plummet, training is fast, and the resulting adapter file is small, often only a few megabytes instead of many gigabytes.