Attention is all you need

The transformer architecture and the attention mechanism that powers every modern LLM.

⏱ 8 min read

In 2017 a Google paper titled Attention Is All You Need introduced the transformer. It threw away recurrence entirely and let every token look directly at every other token. Nearly every frontier model since, from GPT to Claude to Gemini, is built on it.

🔦 The intuition

In "The animal didn't cross the street because it was too tired", what does "it" refer to?

Attention lets the model, while processing "it", look back and put most of its focus on "animal".

Where does “it” look?The2%animal46%didn't3%cross4%the1%street14%because5%it12%was4%too2%tired7%
Attention weights for the token "it": the model attends most strongly to "animal". Darker = more attention.

🔑 Queries, keys and values

Each token produces three vectors: a query (what am I looking for?), a key (what do I contain?), and a value (what will I share?).

Score = query · key for every pair, scaled and softmaxed into weights. Output = weighted sum of values.

Attention(Q,K,V) = softmax(QKᵀ / √d) · V

🧩 Quick quiz

In attention, what is the softmax applied to?

🐙 Multi-head attention

Run several attention operations in parallel, each with its own Q/K/V projections.

Different heads can specialise: one tracks grammar, another tracks which noun a pronoun refers to, another tracks position.

🧱 The transformer block

Each block = multi-head attention → feed-forward network (an MLP applied to each token), with residual connections and layer normalisation around both.

Stack dozens of blocks. Add positional encodings so the model knows word order, since attention itself is order-blind.

🎭 Decoder-only & causal masking

GPT-style LLMs use only the decoder half, with a causal mask: each token may only attend to earlier tokens, never peek at the future.

That's what makes next-token prediction a fair game during training.

🧩 Quick quiz

Why do transformers need positional encodings?

✨ Before you drift off

  • Attention lets each token pull information from any other token.
  • softmax(QKᵀ/√d)·V: score, normalise, mix.
  • Multi-head attention + MLP + residuals + norms = a transformer block.
  • LLMs are decoder-only transformers with causal masking.

📚 Go deeper (free & open)