Fine-tuning on a budget: LoRA & quantisation

How to customise a billion-parameter model on a single GPU.

⏱ 6 min read

Fully fine-tuning a 70-billion-parameter model means updating 70 billion numbers and storing optimiser state for each, which needs a cluster. Parameter-efficient fine-tuning (PEFT) tricks let you get most of the benefit on a single GPU.

W🧊 frozend × d~16.8M numbers+B×Arank r = 8 · 🔥 trainable~65k numbers (0.4%)for one 4096 × 4096 layer
LoRA freezes the big weight matrix W and learns a small low-rank update B·A beside it.

🧊 LoRA: low-rank adaptation

Keep the original weights W frozen. Learn two thin matrices B (d×r) and A (r×d) with a small rank r like 8 or 16.

New output = W·x + B·A·x. For a 4096×4096 layer, that's ~65k trainable numbers at r=8 instead of ~16.8 million.

🔄 Swappable adapters

LoRA adapters are tiny files (megabytes). One base model can host many adapters: one for legal writing, one for your brand voice, one for SQL.

After training, B·A can be merged into W so inference costs nothing extra.

🧩 Quick quiz

In LoRA, which weights are updated during fine-tuning?

🗜️ Quantisation

Store weights in fewer bits: 16-bit → 8-bit → 4-bit. A 4-bit model needs roughly ¼ the memory of 16-bit.

Clever schemes (GPTQ, AWQ, NF4) keep quality surprisingly close to the original.

🤝 QLoRA = both

Quantise the frozen base to 4-bit, then train LoRA adapters on top in higher precision.

This made fine-tuning a 65B model possible on a single 48 GB GPU.

🤔 Should you fine-tune at all?

Try good prompting and RAG first. Fine-tune when you need a consistent style or format, a narrow skill, or a smaller cheaper model that matches a big one on your task.

Fine-tuning is great for behaviour; RAG is usually better for facts.

🧩 Quick quiz

Roughly how much memory does a 4-bit model need compared with the 16-bit version?

✨ Before you drift off

  • LoRA trains tiny low-rank adapters while freezing the base.
  • Quantisation shrinks memory by storing weights in fewer bits.
  • QLoRA combines both for single-GPU fine-tuning.
  • Fine-tune for behaviour, use RAG for knowledge.

📚 Go deeper (free & open)