skip to content

Diffusion and Likelihood Families

Generation by learning to undo noise step by step, plus the exact-likelihood autoregressive and flow families it competes with. Interviewers probe why the sampling chain is the price you pay.

on this pageshow

explore

questions

11

How is a diffusion model trained on one step without simulating the whole noising chain?

level: middleimportance: must knowfreq 72%

answer

  1. the corruption has no parameters
  2. products of (1 - beta) telescope
  3. one Gaussian jump from x_0 to any t
  4. the label is the draw you made

basics

~20 s

The forward corruption is fixed and Gaussian, so any step is one shot: x_t = sqrt(abar_t)*x_0 + sqrt(1-abar_t)*eps, with abar_t the running product of (1-beta). Training draws a random t, builds x_t, and regresses eps.

solid answer

~50 s

The forward process is not learned. Each step scales the sample by `sqrt(1 - beta_t)` and adds Gaussian noise with variance `beta_t`, so composing steps stays Gaussian and the products telescope. Writing `alpha_t = 1 - beta_t` and `abar_t` for the cumulative product, the marginal is `x_t = sqrt(abar_t) * x_0 + sqrt(1 - abar_t) * eps` with `eps` a standard normal draw. That is the whole trick: one training example is a clean sample `x_0`, a step index `t` drawn uniformly, and a fresh `eps`. You build `x_t` directly, feed it to the network along with `t`, and minimise `||eps - eps_theta(x_t, t)||^2`. No chain is walked, no reverse pass is simulated, and the label is free because you drew it yourself. Because `abar_t` shrinks toward zero, the same formula gives a nearly clean sample at small `t` and near-pure noise at the end.

code

python · 21 lines
python
import math, random

T = 1000
betas = [1e-4 + (0.02 - 1e-4) * t / (T - 1) for t in range(T)]
abar, running = [], 1.0
for b in betas:
    running *= (1.0 - b)
    abar.append(running)

def forward_sample(x0, t):
    eps = random.gauss(0.0, 1.0)          # this draw is the regression target
    return math.sqrt(abar[t]) * x0 + math.sqrt(1.0 - abar[t]) * eps, eps

x0 = 1.7                                  # one scalar stand-in for a data point
for t in (0, 300, 600, 999):
    xt, eps = forward_sample(x0, t)
    print(t,
          "signal", round(math.sqrt(abar[t]), 3),
          "noise", round(math.sqrt(1.0 - abar[t]), 3),
          "x_t", round(xt, 3),
          "target", round(eps, 3))

go deeper

for a junior

Be ready to state that the noising direction is fixed and hand-specified while only the denoiser is trained, and that the training label is the noise the code itself just drew.

for a middle

You are expected to write the closed-form marginal from memory, explain where the cumulative product of (1 - beta) comes from, and describe the five-line training step including the uniform step draw.

for a senior

Show the judgment side: what the averaged loss hides, why high-noise steps score better than low-noise ones, and why sampling rather than the loss curve is the evaluation you actually trust.

for a principal

Own the framing that this objective buys a stationary regression target in place of adversarial dynamics, and be able to argue what that trade costs - many network evaluations per sample instead of one.

## What the forward process is A denoising diffusion model defines a **fixed** corruption process that turns a data sample into noise over `T` steps. Each step is ``` x_t = sqrt(1 - beta_t) * x_(t-1) + sqrt(beta_t) * z, z ~ N(0, I) ``` where `beta_t` is a small positive number from a predefined sequence that grows with `t`. Nothing here has parameters: no weights, no gradients, no learning. The scaling by `sqrt(1 - beta_t)` is what makes the process *variance preserving* - if `x_0` has roughly unit variance per dimension, so does every `x_t`, which keeps the network's inputs on one scale across the whole chain. ## The closed form, and why it matters Composing Gaussian steps gives another Gaussian. Define `alpha_t = 1 - beta_t` and `abar_t = alpha_1 * alpha_2 * ... * alpha_t`. Then the marginal of step `t` given the original sample is ``` q(x_t | x_0) = N( sqrt(abar_t) * x_0 , (1 - abar_t) * I ) ``` which you sample as `x_t = sqrt(abar_t) * x_0 + sqrt(1 - abar_t) * eps` with a single standard-normal draw `eps`. The two coefficients are a signal/noise mix whose squares sum to one. At `t = 0` the mix is essentially all signal. Mid-chain the sample is a visible blend - the coarse structure of `x_0` survives while fine detail is gone. As `abar_t` approaches zero the sample is indistinguishable from a draw from `N(0, I)`, which is the tractable prior that sampling starts from. This closed form is the reason diffusion training is cheap. Without it, producing a training input at step 500 would mean running 500 sequential noising operations. With it, every step index is one line of arithmetic, and the steps are trained in random order rather than in sequence. ## The training loop One optimisation step is: 1. Take a clean example `x_0` from the dataset. 2. Draw a step index `t` uniformly from `1..T`. 3. Draw `eps ~ N(0, I)` and build `x_t = sqrt(abar_t) * x_0 + sqrt(1 - abar_t) * eps`. 4. Predict `eps_theta(x_t, t)` - the network sees the noisy sample **and** the step index, usually injected as a positional-style embedding added inside the blocks so one set of weights serves all noise levels. 5. Minimise the squared error against the `eps` you drew. The label is free and exact, which is what distinguishes this from adversarial training: there is no discriminator, no minimax, and no mode-collapse dynamic - just a regression with a stationary target. ## Why regress the noise rather than the clean sample The two are algebraically interchangeable: given `x_t` and a predicted `eps`, the implied clean estimate is `x0_hat = (x_t - sqrt(1 - abar_t) * eps_hat) / sqrt(abar_t)`. What differs is the implicit weighting across noise levels. The `eps` target is a unit-variance quantity at every `t`, so the loss is naturally comparable across the chain, whereas an `x_0` target is nearly trivial at small `t` and dominated by scale factors at large `t`. The simplified unweighted `eps` objective drops the variational weighting terms and, in practice, this reweighting is what makes the objective favour perceptually important noise levels. ## The score-matching view The same object has a second reading. The **score** of a density is the gradient of its log with respect to the input. For the noised conditional above, ``` grad_{x_t} log q(x_t | x_0) = -(x_t - sqrt(abar_t) * x_0) / (1 - abar_t) = -eps / sqrt(1 - abar_t) ``` so a network that predicts the added noise is, up to the factor `-1/sqrt(1 - abar_t)`, estimating the score of the noise-corrupted data density. This is denoising score matching: you cannot compute the score of the data distribution directly, but you can regress the noise you added, and that regression converges to the score of the smoothed density. It is why diffusion models and score-based generative models are two descriptions of one method, and why the reverse chain can be read as Langevin-style movement up the score field. ## What the loss value does and does not tell you The reported loss is an average over randomly drawn `t`, so most of its variance is the step draw, not model progress; it flattens early and then barely moves while samples keep improving. Per-step behaviour is uneven by construction: at large `t` the input is mostly noise, so echoing it back is nearly correct and the loss is low, while at small `t` the little noise present is hard to separate from real detail and the loss is high. Neither pattern indicates a bug. Judging a diffusion model by its training loss curve is the classic mistake - you evaluate by sampling.

  • What is the connection between predicting the added noise and score matching?
    For the noised conditional, the gradient of the log density with respect to `x_t` is `-eps / sqrt(1 - abar_t)`. So a noise-prediction network is a rescaled estimator of the score of the noise-smoothed data density. That is denoising score matching: the score of the data distribution is unavailable, but the noise you injected is a free, unbiased regression target whose optimum is exactly that score.
  • Why regress the noise rather than the clean sample directly?
    They are interchangeable through `x0_hat = (x_t - sqrt(1 - abar_t) * eps_hat) / sqrt(abar_t)`, so the choice is about loss weighting, not expressiveness. The noise target has unit variance at every step, making errors comparable across the chain, while a clean-sample target is trivially easy at low noise and swamped by scale at high noise. The noise parameterisation implicitly emphasises the noise levels that matter perceptually.
  • Why does the training loss flatten early and tell you so little about sample quality?
    Each reported value averages over a uniformly drawn step index, so most of its variance comes from which noise level was sampled rather than from model progress. Per-step difficulty also differs by construction: high-noise steps are easy, low-noise steps are hard. The number is a surrogate objective, not a perceptual metric, so you evaluate by generating samples, not by watching the curve.

The noising chain is like fading a photograph on a known schedule. Because the schedule is known, you can compute exactly how faded it would be after any number of years and jump straight there, instead of waiting through every year.

saying these in an interview costs you the question

  • Says the forward noising process has learned parameters
  • Thinks each training example must walk all T noising steps
  • Forgets the network is conditioned on the step index
  • Treats the added noise variance as constant across steps
  • Reads a flat or falling loss curve as a verdict on sample quality

context

open as a page

Why can an autoregressive image model train in parallel but not sample in parallel?

level: middleimportance: must knowfreq 55%

basics

~20 s

An autoregressive model factorises the joint into conditionals over an ordering. Training scores all of them in one masked pass because the ground truth is already there; sampling must draw each value before the next, one pass per dimension.

open as a page

Why can DDIM sample a diffusion model in 20 steps when ancestral DDPM sampling needed 1000?

level: middleimportance: must knowfreq 72%

basics

~20 s

DDIM reverses the diffusion with a non-Markovian, noise-free update that matches the same training marginals, so one trained network can be run on any subsequence of timesteps. Because the update is deterministic, dropping steps degrades gracefully instead of breaking.

open as a page

Why did diffusion models move from a linear noise schedule to a cosine one?

level: middleimportance: should knowfreq 55%

basics

~20 s

A linear beta schedule destroys nearly all signal two-thirds of the way through the chain, so the final third of the reverse steps teaches the model almost nothing. A cosine schedule lets signal-to-noise fall gradually to the end.

open as a page

How does classifier-free guidance steer a diffusion model without training a classifier?

level: seniorimportance: should knowfreq 58%

basics

~20 s

One network learns both conditional and unconditional prediction, because the condition is replaced by a learned null token on roughly ten percent of training examples. Sampling evaluates it twice per step and extrapolates along the difference between the two predictions.

open as a page

In a diffusion model, why can't the reverse go from pure noise to data in one step?

level: seniorimportance: should knowfreq 55%

basics

~20 s

The forward jump is closed-form because it conditions on the clean sample. Backwards, many clean samples explain one noisy one, so the reverse distribution is multimodal; a single pass returns only its blurry average. Small steps keep each reverse move Gaussian.

open as a page

Why must a normalizing flow be invertible with a cheap log-determinant Jacobian?

level: seniorimportance: should knowfreq 35%

basics

~20 s

A flow's exact likelihood comes from the change-of-variables formula, which needs the inverse map and the log-determinant of its Jacobian. A general determinant costs cubic time, so flow layers are built to have a triangular Jacobian.

open as a page

In diffusion training, what does v-prediction fix that epsilon-prediction breaks at high noise?

level: seniorimportance: should knowfreq 35%

basics

~20 s

At the noisiest steps the input is almost pure noise, so predicting the added noise is nearly trivial while small errors explode when converted back to an image. v-prediction blends the noise and image targets, staying informative at both ends.

open as a page

In diffusion modelling, when is a fixed Gaussian forward process the wrong fit for your data?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Gaussian corruption assumes continuous, comparably scaled features. It suits smooth vector data such as robot action trajectories, and suits discrete tokens, hard-constrained quantities and heavy-tailed features badly - and every sample costs many network evaluations.

open as a page

When a flow beats a GAN on exact likelihood but its samples look worse, what do you conclude?

level: principalimportance: nice to knowfreq 28%

basics

~10 s

Likelihood and sample quality measure different things. Maximum likelihood minimises a mode-covering divergence that punishes missing data but not wasted mass, so a better likelihood is evidence of density fit, not of better-looking samples.

open as a page

When is distilling a diffusion sampler to four steps worth it over just cutting sampler steps?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Cut steps and change solver first - both are free and reversible. Distillation earns its cost only when a hard latency floor sits below what any training-free sampler reaches, and you accept narrower diversity plus a retraining stage per model version.

open as a page