MLGuerrillaStart with M1 →
Architectures·16 min read·Updated 24 September 2026

Diffusion models

Diffusion is the architecture behind most image and video generators. A diffusion model starts from pure noise and removes a little of it at a time until a picture is left.

A diffusion model generates by removing noise. Training shows it images with a known amount of noise added and asks it to predict that noise. Generation starts from an image of pure noise and applies the model over and over, taking a little noise out each time, until something coherent is left.

The idea comes from Ho et al. (2020), and it's behind most image and video generators you can name, along with a growing number of models that generate other things, from audio to robot actions.

What it is

Noise goes in one direction to train, and comes out the other to generate

The training side is arithmetic with no model in it. Take a real image, pick a step number t between 1 and 1000, and mix in Gaussian noise according to a fixed schedule, so a low t barely changes the picture and a high t leaves nothing. The model's job is to look at that noisy image, together with t, and say what the noise was.

Generation runs the reverse. Start from pure noise, ask the model what noise it sees, subtract a portion of it, and repeat. Each pass is one step, and a modern model needs between 1 and 50 of them.

A figure titled "Training destroys an image with noise, and generation runs the other way", in two rows. The top row, labeled "Forward: add noise on a fixed schedule, arithmetic only, with no model involved", shows the same street photo six times at t equals 0, 100, 250, 500, 750 and 999, with the signal remaining falling from 100% to 94.6%, 72.2%, 27.9%, 5.7% and 0.6%, so the last two images look like colored static. The bottom row, labeled "Reverse: a trained model removes noise, SD-Turbo, same prompt and seed, on a CPU", starts from a pure noise square and shows three generated photos of a red bicycle against a brick wall, at 1 step taking 31 seconds, 2 steps taking 45 seconds and 4 steps taking 75 seconds. The prompt was "a red bicycle parked against a brick wall, photograph". The caption says the model was trained to predict the noise in a picture like the top row, and generation applies that guess again and again.
The bicycle images are a real run on a laptop CPU. On a GPU the same model produces these in well under a second.

The percentages in the top row are the schedule itself. At t = 500, the original image contributes 27.9% of the signal and noise contributes the rest, and by t = 999 almost nothing of the photo is left. Training samples a t at random for every image, so the model sees every level of destruction and learns to undo each one.

Where it shows up

It shows up wherever something continuous gets generated

Images from a prompt

This is the job most people mean when they say image generation. A text encoder turns the prompt into vectors, the denoiser is conditioned on them, and the output is an image that matches the words as closely as the guidance setting demands.

Editing an existing image

Inpainting and image-to-image both start from a partly noised version of a real image, where generation from scratch starts from pure noise. The model fills in what you masked, or redraws the whole thing at a noise level you choose, which is how "keep the composition, change the style" works.

Video

Video models diffuse a sequence of frames together so the frames stay consistent, which is why generated clips are short and expensive next to still images.

Audio and other continuous signals

The same recipe generates spectrograms for speech and music. It also generates robot actions, where a policy denoises a short trajectory in place of a picture.

Text, in a few new models

The LLM article mentions diffusion language models, which denoise all the tokens in parallel over a few steps. They work differently from the token-by-token loop that decoder-only models run.

How it works

Predict the noise, subtract some of it, and repeat

Training predicts the noise that was added

For each training image the model gets the noised version and the step number, and it predicts the noise that was added. The loss is the difference between the predicted noise and the real noise, which is a plain regression problem. That's the whole objective, and everything else is machinery around it.

Sampling is the same model applied many times

A sampler, also called a scheduler, decides how much noise to remove at each step and how to move from one step to the next. DDIM (Song et al., 2020) showed that a deterministic sampler can skip most of the 1000 training steps, which is why generation with 20 to 50 steps became normal.

Guidance decides how closely it follows the prompt

Classifier-free guidance (Ho and Salimans, 2022) runs the model twice at each step, once with the prompt and once without, then pushes the prediction away from the unconditional one. The guidance scale controls how hard it pushes. Low values wander off the prompt, and high values produce oversaturated images that follow the prompt rigidly. Defaults differ by model. As of September 2026, the diffusers pipeline default is 5.0 for SDXL and 7.0 for Stable Diffusion 3, and Stability's own example for Stable Diffusion 3.5 Large uses 3.5. Guidance also doubles the compute per step, which is why some fast models are trained to work without it. SD-Turbo is one of them, and its model card says it "does not make use of guidance_scale".

Latent diffusion moves the work into a smaller space

Denoising full-resolution pixels is expensive. Latent diffusion (Rombach et al., 2021), the design behind Stable Diffusion, first compresses the image with a variational autoencoder, runs the whole diffusion process in that smaller space, and decodes only at the end. A 512 × 512 image becomes a 64 × 64 latent, so every step is far cheaper. This is why the variational autoencoder entry on this index belongs to image generation.

The denoiser used to be a U-Net and is now often a transformer

Early diffusion models used a U-Net, a convolutional network built for segmentation, with the prompt fed in through cross-attention. Diffusion transformers (Peebles and Xie, 2022) replaced it with a transformer over latent patches and found that it scales better, and current large models follow that path. Stable Diffusion 3 (Esser et al., 2024) combined a transformer backbone with rectified flow, a training recipe from the flow matching line that learns a straighter path from noise to image and needs fewer steps.

Distillation cuts the steps to a handful

A distilled model is trained to match, in one or two steps, what a slower model produces in fifty. Adversarial Diffusion Distillation (Sauer et al., 2023) is the method behind SD-Turbo, the model in the figure, which is why a 1-step image is usable at all. The figure's own timings show the cost of extra steps, from 31 seconds at 1 step to 75 seconds at 4 on a CPU.

Versions

The models you'll see, as of September 2026

Image generation splits into closed models you call through an API and open weights you can run yourself. The frontier image models entry on this index tracks the closed ones, and these are the open weights people build on.

Open image models on Hugging Face
ModelFromPublishedLicense
black-forest-labs/FLUX.2-devBlack Forest LabsNovember 2025Custom, non-commercial for the dev weights
stabilityai/stable-diffusion-3.5-largeStability AIOctober 2024Stability community license
Qwen/Qwen-ImageAlibabaAugust 2025Apache 2.0

Licenses matter more here than in most model families, because several popular checkpoints allow research and personal use while charging for commercial use. Check the license file of the exact checkpoint before shipping, since two models from the same lab often differ.

Choosing

Pick diffusion for images, and know what it competes with

Which generator to reach for
The jobReach forWhy
Images from a prompt, run on your own hardwareAn open diffusion model, like FLUX or Stable DiffusionWeights you can run, fine-tune and steer with adapters
The best image quality with no infrastructureA closed image model through an APIThe frontier models lead on prompt following and text rendering
Images that must contain readable text or follow a long instructionA current frontier model, closed or the newest open oneOlder diffusion models garble text, and this is where recent work has concentrated
The same image many times, as fast as possibleA distilled few-step model1 to 4 steps in place of 30, at some cost in detail
A fast image with a small model and no prompt fidelity needsA GANOne forward pass with no iteration, and no text conditioning by default
Text or codeAn LLMDiffusion generates continuous signals, and language is discrete

Try it

How to try it

The diffusers library runs most open models with the same few lines. This is the code that produced the bottom row of the figure.

generate.py — a few-step image on a CPU
import torch
from diffusers import AutoPipelineForText2Image

pipe = AutoPipelineForText2Image.from_pretrained("stabilityai/sd-turbo",
                                                 torch_dtype=torch.float32)

image = pipe(prompt="a red bicycle parked against a brick wall, photograph",
             num_inference_steps=4,      # SD-Turbo is distilled for 1 to 4
             guidance_scale=0.0,         # and trained to need no guidance
             generator=torch.Generator().manual_seed(7)).images[0]

image.save("bicycle.jpg")

Keep the generator seed fixed while you change one setting at a time, since the same seed with the same prompt gives the same image and makes the effect of a change visible.

SD-Turbo can show what extra steps cost, but it can't show what guidance does, because it was distilled to run in 1 to 4 steps with guidance switched off. Sweeping its guidance scale or running it for 30 steps shows settings it was never trained for. A guidance sweep needs a model that was trained with classifier-free guidance, like SDXL or Stable Diffusion 3.5, at the step count its card recommends. The two prompts below keep those experiments apart.

Ask your AI coding tool, for the speed demo

Write a script that shows me what extra denoising steps cost with a distilled model. Using diffusers with stabilityai/sd-turbo, generate one prompt I pass in at 1, 2, 3 and 4 inference steps, with guidance_scale=0.0 and the same seed every time, since the model was trained for exactly that range and without guidance. Run one warm-up generation first and then time each run. Save the four images side by side as a PNG labeled with the step count and the wall-clock time. Print whether the run used a CPU or a GPU.

Ask your AI coding tool, for the guidance sweep

Write a script that shows me what the guidance scale does. Using diffusers with stabilityai/stable-diffusion-xl-base-1.0 (or stabilityai/stable-diffusion-3.5-large if I have a GPU with enough memory), generate one prompt I pass in at guidance scales 1, 3, 5, 7 and 10, keeping the step count fixed at the model's usual setting (50 for SDXL, which is the diffusers default, and 28 for SD 3.5 Large, which is Stability's example) and the seed fixed in every image. Save the images in a row as a PNG labeled with the guidance scale. Then repeat the row with four different seeds at the best-looking scale, so I can see how much of the result comes from the seed and how much from the setting.

Limits

What it can't do

  • It's slow compared with a single forward pass. Every step is a full run of the denoiser, and guidance doubles that, so a 30-step image with guidance is 60 model calls.
  • The prompt is a suggestion. Counts and spatial relations come out wrong often, and so do negations, and the usual workarounds are regenerating with a new seed or controlling the layout with an adapter.
  • Small details break. Hands, small faces and fine text are where the errors concentrate, which is why models are often judged on exactly those.
  • Two runs with the same seed and different libraries won't match. Reproducibility depends on the sampler and the step count, and on the library version and hardware too.
  • It generates, and it can't check. Nothing in the process knows whether the thing it drew is safe or legal to use, so filtering and review belong in your pipeline.
  • Licensing is a real constraint. Several of the strongest open checkpoints are free for research and paid for commercial use, and the weights of the strongest closed models aren't available at all.

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.