What it is
Noise goes in one direction to train, and comes out the other to generate
The training side is arithmetic with no model in it. Take a real image, pick a step number t between 1 and 1000, and mix in Gaussian noise according to a fixed schedule, so a low t barely changes the picture and a high t leaves nothing. The model's job is to look at that noisy image, together with t, and say what the noise was.
Generation runs the reverse. Start from pure noise, ask the model what noise it sees, subtract a portion of it, and repeat. Each pass is one step, and a modern model needs between 1 and 50 of them.

The percentages in the top row are the schedule itself. At t = 500, the original image contributes 27.9% of the signal and noise contributes the rest, and by t = 999 almost nothing of the photo is left. Training samples a t at random for every image, so the model sees every level of destruction and learns to undo each one.
Where it shows up
It shows up wherever something continuous gets generated
Images from a prompt
This is the job most people mean when they say image generation. A text encoder turns the prompt into vectors, the denoiser is conditioned on them, and the output is an image that matches the words as closely as the guidance setting demands.
Editing an existing image
Inpainting and image-to-image both start from a partly noised version of a real image, where generation from scratch starts from pure noise. The model fills in what you masked, or redraws the whole thing at a noise level you choose, which is how "keep the composition, change the style" works.
Video
Video models diffuse a sequence of frames together so the frames stay consistent, which is why generated clips are short and expensive next to still images.
Audio and other continuous signals
The same recipe generates spectrograms for speech and music. It also generates robot actions, where a policy denoises a short trajectory in place of a picture.
Text, in a few new models
The LLM article mentions diffusion language models, which denoise all the tokens in parallel over a few steps. They work differently from the token-by-token loop that decoder-only models run.
How it works
Predict the noise, subtract some of it, and repeat
Training predicts the noise that was added
For each training image the model gets the noised version and the step number, and it predicts the noise that was added. The loss is the difference between the predicted noise and the real noise, which is a plain regression problem. That's the whole objective, and everything else is machinery around it.
Sampling is the same model applied many times
A sampler, also called a scheduler, decides how much noise to remove at each step and how to move from one step to the next. DDIM (Song et al., 2020) showed that a deterministic sampler can skip most of the 1000 training steps, which is why generation with 20 to 50 steps became normal.
Guidance decides how closely it follows the prompt
Classifier-free guidance (Ho and Salimans, 2022) runs the model twice at each step, once with the prompt and once without, then pushes the prediction away from the unconditional one. The guidance scale controls how hard it pushes. Low values wander off the prompt, and high values produce oversaturated images that follow the prompt rigidly. Defaults differ by model. As of September 2026, the diffusers pipeline default is 5.0 for SDXL and 7.0 for Stable Diffusion 3, and Stability's own example for Stable Diffusion 3.5 Large uses 3.5. Guidance also doubles the compute per step, which is why some fast models are trained to work without it. SD-Turbo is one of them, and its model card says it "does not make use of guidance_scale".
Latent diffusion moves the work into a smaller space
Denoising full-resolution pixels is expensive. Latent diffusion (Rombach et al., 2021), the design behind Stable Diffusion, first compresses the image with a variational autoencoder, runs the whole diffusion process in that smaller space, and decodes only at the end. A 512 × 512 image becomes a 64 × 64 latent, so every step is far cheaper. This is why the variational autoencoder entry on this index belongs to image generation.
The denoiser used to be a U-Net and is now often a transformer
Early diffusion models used a U-Net, a convolutional network built for segmentation, with the prompt fed in through cross-attention. Diffusion transformers (Peebles and Xie, 2022) replaced it with a transformer over latent patches and found that it scales better, and current large models follow that path. Stable Diffusion 3 (Esser et al., 2024) combined a transformer backbone with rectified flow, a training recipe from the flow matching line that learns a straighter path from noise to image and needs fewer steps.
Distillation cuts the steps to a handful
A distilled model is trained to match, in one or two steps, what a slower model produces in fifty. Adversarial Diffusion Distillation (Sauer et al., 2023) is the method behind SD-Turbo, the model in the figure, which is why a 1-step image is usable at all. The figure's own timings show the cost of extra steps, from 31 seconds at 1 step to 75 seconds at 4 on a CPU.
Versions
The models you'll see, as of September 2026
Image generation splits into closed models you call through an API and open weights you can run yourself. The frontier image models entry on this index tracks the closed ones, and these are the open weights people build on.
| Model | From | Published | License |
|---|---|---|---|
black-forest-labs/FLUX.2-dev | Black Forest Labs | November 2025 | Custom, non-commercial for the dev weights |
stabilityai/stable-diffusion-3.5-large | Stability AI | October 2024 | Stability community license |
Qwen/Qwen-Image | Alibaba | August 2025 | Apache 2.0 |
Licenses matter more here than in most model families, because several popular checkpoints allow research and personal use while charging for commercial use. Check the license file of the exact checkpoint before shipping, since two models from the same lab often differ.
Choosing
Pick diffusion for images, and know what it competes with
| The job | Reach for | Why |
|---|---|---|
| Images from a prompt, run on your own hardware | An open diffusion model, like FLUX or Stable Diffusion | Weights you can run, fine-tune and steer with adapters |
| The best image quality with no infrastructure | A closed image model through an API | The frontier models lead on prompt following and text rendering |
| Images that must contain readable text or follow a long instruction | A current frontier model, closed or the newest open one | Older diffusion models garble text, and this is where recent work has concentrated |
| The same image many times, as fast as possible | A distilled few-step model | 1 to 4 steps in place of 30, at some cost in detail |
| A fast image with a small model and no prompt fidelity needs | A GAN | One forward pass with no iteration, and no text conditioning by default |
| Text or code | An LLM | Diffusion generates continuous signals, and language is discrete |
Try it
How to try it
The diffusers library runs most open models with the same few lines. This is the code that produced the bottom row of the figure.
import torch
from diffusers import AutoPipelineForText2Image
pipe = AutoPipelineForText2Image.from_pretrained("stabilityai/sd-turbo",
torch_dtype=torch.float32)
image = pipe(prompt="a red bicycle parked against a brick wall, photograph",
num_inference_steps=4, # SD-Turbo is distilled for 1 to 4
guidance_scale=0.0, # and trained to need no guidance
generator=torch.Generator().manual_seed(7)).images[0]
image.save("bicycle.jpg")Keep the generator seed fixed while you change one setting at a time, since the same seed with the same prompt gives the same image and makes the effect of a change visible.
SD-Turbo can show what extra steps cost, but it can't show what guidance does, because it was distilled to run in 1 to 4 steps with guidance switched off. Sweeping its guidance scale or running it for 30 steps shows settings it was never trained for. A guidance sweep needs a model that was trained with classifier-free guidance, like SDXL or Stable Diffusion 3.5, at the step count its card recommends. The two prompts below keep those experiments apart.
Write a script that shows me what extra denoising steps cost with a distilled model. Using diffusers with stabilityai/sd-turbo, generate one prompt I pass in at 1, 2, 3 and 4 inference steps, with guidance_scale=0.0 and the same seed every time, since the model was trained for exactly that range and without guidance. Run one warm-up generation first and then time each run. Save the four images side by side as a PNG labeled with the step count and the wall-clock time. Print whether the run used a CPU or a GPU.
Write a script that shows me what the guidance scale does. Using diffusers with stabilityai/stable-diffusion-xl-base-1.0 (or stabilityai/stable-diffusion-3.5-large if I have a GPU with enough memory), generate one prompt I pass in at guidance scales 1, 3, 5, 7 and 10, keeping the step count fixed at the model's usual setting (50 for SDXL, which is the diffusers default, and 28 for SD 3.5 Large, which is Stability's example) and the seed fixed in every image. Save the images in a row as a PNG labeled with the guidance scale. Then repeat the row with four different seeds at the best-looking scale, so I can see how much of the result comes from the seed and how much from the setting.
Limits
What it can't do
- It's slow compared with a single forward pass. Every step is a full run of the denoiser, and guidance doubles that, so a 30-step image with guidance is 60 model calls.
- The prompt is a suggestion. Counts and spatial relations come out wrong often, and so do negations, and the usual workarounds are regenerating with a new seed or controlling the layout with an adapter.
- Small details break. Hands, small faces and fine text are where the errors concentrate, which is why models are often judged on exactly those.
- Two runs with the same seed and different libraries won't match. Reproducibility depends on the sampler and the step count, and on the library version and hardware too.
- It generates, and it can't check. Nothing in the process knows whether the thing it drew is safe or legal to use, so filtering and review belong in your pipeline.
- Licensing is a real constraint. Several of the strongest open checkpoints are free for research and paid for commercial use, and the weights of the strongest closed models aren't available at all.
Go deeper
Ho et al. (2020): Denoising Diffusion Probabilistic Models · Song et al. (2020): Denoising Diffusion Implicit Models (DDIM) · Rombach et al. (2021): High-Resolution Image Synthesis with Latent Diffusion Models · Ho and Salimans (2022): Classifier-Free Diffusion Guidance · Peebles and Xie (2022): Scalable Diffusion Models with Transformers (DiT) · Esser et al. (2024): Scaling Rectified Flow Transformers (Stable Diffusion 3) · Sauer et al. (2023): Adversarial Diffusion Distillation (SD-Turbo)
Coming soon
New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.
