Diffusion: art from noise
How Stable Diffusion and friends turn static into pictures, one denoising step at a time.
Image generators like Stable Diffusion, DALL·E and Midjourney are built on a beautifully odd idea: teach a network to remove noise, then start from pure noise and let it keep cleaning until a picture appears.
➡️ Forward: destroy
Take a real image and add a little Gaussian noise. Repeat for, say, 1,000 steps until it's pure static.
This direction needs no learning; it's just maths.
⬅️ Reverse: create
Train a network (often a U-Net or a transformer) to look at a noisy image and predict the noise that was added.
Subtract the predicted noise, step by step, and an image emerges from static.
What is a diffusion model usually trained to predict?
📝 Text-to-image
To follow a prompt, the denoiser is conditioned on a text embedding (e.g. from CLIP or a T5 encoder) via cross-attention.
Classifier-free guidance pushes outputs harder towards the prompt; the "guidance scale" slider controls how strongly.
🗜️ Latent diffusion
Denoising full-resolution pixels is expensive. Stable Diffusion works in a compressed latent space from an autoencoder, then decodes to pixels at the end.
Same idea, far cheaper, which is why it runs on consumer GPUs.
Why does Stable Diffusion work in latent space?
✨ Before you drift off
- Forward process: add noise. Reverse process: learn to remove it.
- Models predict noise, then subtract it step by step.
- Text conditioning + guidance steer the image.
- Latent diffusion makes it efficient.
📚 Go deeper (free & open)
- Denoising Diffusion Probabilistic Models ↗ · Ho, Jain & Abbeel, 2020
- What are Diffusion Models? ↗ · Lilian Weng
- The Illustrated Stable Diffusion ↗ · Jay Alammar
- Hugging Face Diffusion Models Course ↗ · Hugging Face · Apache 2.0