Fine-tuning on a budget: LoRA & quantisation
How to customise a billion-parameter model on a single GPU.
Fully fine-tuning a 70-billion-parameter model means updating 70 billion numbers and storing optimiser state for each, which needs a cluster. Parameter-efficient fine-tuning (PEFT) tricks let you get most of the benefit on a single GPU.
🧊 LoRA: low-rank adaptation
Keep the original weights W frozen. Learn two thin matrices B (d×r) and A (r×d) with a small rank r like 8 or 16.
New output = W·x + B·A·x. For a 4096×4096 layer, that's ~65k trainable numbers at r=8 instead of ~16.8 million.
🔄 Swappable adapters
LoRA adapters are tiny files (megabytes). One base model can host many adapters: one for legal writing, one for your brand voice, one for SQL.
After training, B·A can be merged into W so inference costs nothing extra.
In LoRA, which weights are updated during fine-tuning?
🗜️ Quantisation
Store weights in fewer bits: 16-bit → 8-bit → 4-bit. A 4-bit model needs roughly ¼ the memory of 16-bit.
Clever schemes (GPTQ, AWQ, NF4) keep quality surprisingly close to the original.
🤝 QLoRA = both
Quantise the frozen base to 4-bit, then train LoRA adapters on top in higher precision.
This made fine-tuning a 65B model possible on a single 48 GB GPU.
🤔 Should you fine-tune at all?
Try good prompting and RAG first. Fine-tune when you need a consistent style or format, a narrow skill, or a smaller cheaper model that matches a big one on your task.
Fine-tuning is great for behaviour; RAG is usually better for facts.
Roughly how much memory does a 4-bit model need compared with the 16-bit version?
✨ Before you drift off
- LoRA trains tiny low-rank adapters while freezing the base.
- Quantisation shrinks memory by storing weights in fewer bits.
- QLoRA combines both for single-GPU fine-tuning.
- Fine-tune for behaviour, use RAG for knowledge.
📚 Go deeper (free & open)
- LoRA: Low-Rank Adaptation of Large Language Models ↗ · Hu et al., 2021
- QLoRA: Efficient Finetuning of Quantized LLMs ↗ · Dettmers et al., 2023
- PEFT library documentation ↗ · Hugging Face · Apache 2.0