Attention is all you need
The transformer architecture and the attention mechanism that powers every modern LLM.
In 2017 a Google paper titled Attention Is All You Need introduced the transformer. It threw away recurrence entirely and let every token look directly at every other token. Nearly every frontier model since, from GPT to Claude to Gemini, is built on it.
🔦 The intuition
In "The animal didn't cross the street because it was too tired", what does "it" refer to?
Attention lets the model, while processing "it", look back and put most of its focus on "animal".
🔑 Queries, keys and values
Each token produces three vectors: a query (what am I looking for?), a key (what do I contain?), and a value (what will I share?).
Score = query · key for every pair, scaled and softmaxed into weights. Output = weighted sum of values.
Attention(Q,K,V) = softmax(QKᵀ / √d) · V
In attention, what is the softmax applied to?
🐙 Multi-head attention
Run several attention operations in parallel, each with its own Q/K/V projections.
Different heads can specialise: one tracks grammar, another tracks which noun a pronoun refers to, another tracks position.
🧱 The transformer block
Each block = multi-head attention → feed-forward network (an MLP applied to each token), with residual connections and layer normalisation around both.
Stack dozens of blocks. Add positional encodings so the model knows word order, since attention itself is order-blind.
🎭 Decoder-only & causal masking
GPT-style LLMs use only the decoder half, with a causal mask: each token may only attend to earlier tokens, never peek at the future.
That's what makes next-token prediction a fair game during training.
Why do transformers need positional encodings?
✨ Before you drift off
- Attention lets each token pull information from any other token.
- softmax(QKᵀ/√d)·V: score, normalise, mix.
- Multi-head attention + MLP + residuals + norms = a transformer block.
- LLMs are decoder-only transformers with causal masking.
📚 Go deeper (free & open)
- Attention Is All You Need ↗ · Vaswani et al., 2017
- The Illustrated Transformer ↗ · Jay Alammar
- Transformer Explainer (interactive) ↗ · Georgia Tech Polo Club · MIT licensed, runs GPT-2 in your browser
- Let's build GPT: from scratch, in code ↗ · Andrej Karpathy