Diffusion: art from noise

How Stable Diffusion and friends turn static into pictures, one denoising step at a time.

⏱ 6 min read

Image generators like Stable Diffusion, DALL·E and Midjourney are built on a beautifully odd idea: teach a network to remove noise, then start from pure noise and let it keep cleaning until a picture appears.

forward: add noise (no learning needed)imagestep 1step 2step 3pure noisereverse: the model predicts & removes noisestart from random noise → get a brand-new image
Forward process adds noise step by step. The model learns the reverse: predicting and removing noise to recover an image.

➡️ Forward: destroy

Take a real image and add a little Gaussian noise. Repeat for, say, 1,000 steps until it's pure static.

This direction needs no learning; it's just maths.

⬅️ Reverse: create

Train a network (often a U-Net or a transformer) to look at a noisy image and predict the noise that was added.

Subtract the predicted noise, step by step, and an image emerges from static.

🧩 Quick quiz

What is a diffusion model usually trained to predict?

📝 Text-to-image

To follow a prompt, the denoiser is conditioned on a text embedding (e.g. from CLIP or a T5 encoder) via cross-attention.

Classifier-free guidance pushes outputs harder towards the prompt; the "guidance scale" slider controls how strongly.

🗜️ Latent diffusion

Denoising full-resolution pixels is expensive. Stable Diffusion works in a compressed latent space from an autoencoder, then decodes to pixels at the end.

Same idea, far cheaper, which is why it runs on consumer GPUs.

🧩 Quick quiz

Why does Stable Diffusion work in latent space?

✨ Before you drift off

  • Forward process: add noise. Reverse process: learn to remove it.
  • Models predict noise, then subtract it step by step.
  • Text conditioning + guidance steer the image.
  • Latent diffusion makes it efficient.

📚 Go deeper (free & open)