How Diffusion Models Work
A step-by-step, visual explainer of forward noise and reverse denoising in modern diffusion models.
Diffusion Models
Diffusion models were introduced as generative probabilistic frameworks that gradually add and then remove noise. By learning to reverse the diffusion, they synthesize new samples that follow the data distribution.
Forward Process
The forward process applies Gaussian noise over a schedule of discrete time steps:Sampling uses the reparameterization trick:
so a single diffusion step is
With the definitions
we get the closed form:
Reverse Process
The learned model predicts the reverse transition:Predicting the model mean is equivalent to predicting the noise:
Using this estimate, the posterior mean is:
The reverse step samples:
starting from pure noise.
Neural network view
Under the hood, a neural network predicts the noise at each timestep. The animation below shows a dense network with signals pulsing from input to output.
Self-attention over pixels
Modern diffusion U-Nets often mix convolutions with self-attention so distant pixels can influence each other. Click any pixel below to make it the query; the white lines show attention weights fading with distance, controlled by the σ slider.
U-Net backbone
Diffusion models commonly use a U-Net: an encoder that downsamples to a bottleneck, then a decoder that upsamples while fusing skip connections. Slide to see the forward (left) and reverse (right) halves light up.
Training dynamics
Diffusion models are optimized with gradient descent. The plot below shows steps on a simple quadratic: small learning rates creep toward the minimum; larger rates move faster but can overshoot.
Single gate intuition
A neural network is built from simple gates. Here is one ReLU-style gate with two inputs and one output—signals flow left to right in monochrome.