How LLMs are trained

Pre-training, supervised fine-tuning, and learning from feedback: the three-stage recipe.

⏱ 7 min read

A chat assistant isn't trained in one go. It's built in stages, each with a different goal and very different kinds of data.

data size1 · Pre-trainingnext-token prediction, web scaletrillions of tokens2 · Supervised fine-tuninglearn from ideal example answers~100k examples3 · Preference tuningRLHF / DPO / AI feedbackcomparisons🤖 Assistantbase model is madein stage 1; stages2–3 shape behaviour
The typical recipe: a huge pre-training run, then much smaller but carefully curated post-training stages.

1️⃣ Pre-training

Next-token prediction on trillions of tokens of web pages, books, code and more.

Takes thousands of GPUs for weeks or months. Produces a base model: vast knowledge, but it just continues text. Ask it a question and it might reply with more questions.

2️⃣ Supervised fine-tuning (SFT)

Train on high-quality example conversations: a prompt and an ideal response, written or curated by people (and increasingly by other models).

The model learns the format and style of being a helpful assistant.

3️⃣ Learning from feedback

RLHF: humans compare pairs of answers; a reward model learns to predict their preferences; RL (classically PPO) optimises the LLM against it.

DPO skips the separate reward model and learns directly from preference pairs.

Constitutional AI / RLAIF uses a written set of principles and AI feedback to scale this up.

🧩 Quick quiz

What does a base model (only pre-trained) typically do?

🧠 Reasoning models

Newer models are also trained with RL on tasks with checkable answers (maths, code) to produce long chains of thought before answering.

Spending more compute at inference time (thinking longer) turns out to boost accuracy on hard problems.

🧹 Data is still king

Pre-training data is filtered, de-duplicated and mixed carefully (how much code? how many languages?).

Open datasets like FineWeb and The Pile let researchers study this in public.

🧩 Quick quiz

In RLHF, what does the reward model learn?

✨ Before you drift off

  • Pre-training → base model with broad knowledge.
  • SFT → assistant format and style.
  • RLHF / DPO / RLAIF → aligned with preferences.
  • RL on verifiable tasks produces stronger reasoning.

📚 Go deeper (free & open)