How does classifier-free guidance steer a diffusion model without training a classifier?
answer
- no second model is involved
- the condition is sometimes hidden during training
- one learned token means 'nothing given'
- two predictions, combined along their difference
basics
~20 sOne network learns both conditional and unconditional prediction, because the condition is replaced by a learned null token on roughly ten percent of training examples. Sampling evaluates it twice per step and extrapolates along the difference between the two predictions.
solid answer
~50 sDuring training, the conditioning input - a text embedding, a class label, a goal vector - is randomly replaced by a learned null token on a small fraction of examples, commonly around ten percent. The same weights therefore learn both `eps_theta(x_t, t, c)` and `eps_theta(x_t, t, null)`. At sampling, each reverse step evaluates the network twice and mixes the predictions: `eps_hat = eps_uncond + w * (eps_cond - eps_uncond)`. The difference vector points from the generic prediction toward the conditioned one, so `w = 1` reproduces ordinary conditional sampling and `w > 1` extrapolates past it, pushing the trajectory further into the region the condition specifies. The earlier approach, classifier guidance, needed a separate classifier trained on noisy inputs at every noise level so its label gradient could be added to the score. Classifier-free guidance removes that second model entirely - the price is two network evaluations per step instead of one.
go deeper
Recall that conditioning is fed to the denoiser as an extra input alongside the noisy sample and the step index, and that guidance strengthens how closely samples follow that input.
Explain the two halves of the mechanism: randomly replacing the condition with a null token during training, and combining conditional with unconditional predictions at each reverse step.
Show you know why it works - the difference between the two predictions is an implicit label gradient - and account for the real costs: doubled evaluations per step and a dropout rate that must be chosen.
Be ready to weigh conditioning strategies as an architecture decision: an implicit gradient from one generator against an explicit auxiliary model, and what each implies for serving cost, retraining coupling and the kinds of conditions you can support.
## The problem it solves A conditional diffusion model can be trained straightforwardly: feed the condition `c` alongside the noisy input and the step index, and the network learns `eps_theta(x_t, t, c)`. Left alone, such models often follow the condition loosely - the conditioning signal is one input among many, and nothing in the objective forces the samples to be strongly characteristic of `c`. The first fix was **classifier guidance**. Train a separate classifier that accepts *noisy* inputs at every noise level, take the gradient of `log p(c | x_t)` with respect to `x_t`, and add it to the score estimate the diffusion model provides. It works, but it demands a second model trained on the corrupted-input distribution at all noise levels - an awkward artefact to build, maintain and keep in sync with the generator, and one that only supports conditions expressible as classifier outputs. ## The construction Classifier-free guidance obtains the same steering from a single network by making it serve two roles. **Training.** Keep the ordinary conditional objective, but with some probability - commonly about ten percent - replace `c` with a learned null embedding standing for 'no condition'. Because the step index is drawn uniformly and the dropout is independent per example, the same weights end up estimating both the conditional noise and the unconditional noise across the whole chain. There is no second network, no second optimiser, no extra loss term. **Sampling.** At each reverse step, evaluate the network twice - once with `c`, once with the null token - and combine: ``` eps_hat = eps_uncond + w * (eps_cond - eps_uncond) ``` Equivalently `eps_hat = (1 - w) * eps_uncond + w * eps_cond`. Setting `w = 0` gives unconditional generation; `w = 1` gives ordinary conditional generation; `w > 1` extrapolates beyond the conditional prediction along the direction that distinguishes conditioned from unconditioned behaviour. The combined `eps_hat` is then fed into the usual reverse-step arithmetic in place of the raw prediction. ## Why the difference vector is the right direction Read through the score lens. The noise prediction is a rescaled score estimate, so `eps_cond - eps_uncond` is proportional to the difference between the score of the conditional noised density and the score of the unconditional one. By Bayes' rule that difference is proportional to the gradient of `log p(c | x_t)` - the very quantity classifier guidance obtained from an external classifier. The two conditional evaluations of one generator therefore recover an *implicit* classifier gradient without anyone ever training a classifier. That equivalence is the theoretical content of the method and the answer interviewers are listening for. ## Practical consequences - **Cost.** Two evaluations per reverse step. In practice the conditional and unconditional inputs are stacked into a single batch, so the wall-clock hit is closer to a wider batch than to a doubled step count, but the arithmetic is genuinely doubled. - **The null token is learned, not zero.** It is an embedding trained alongside everything else; zeroing the conditioning inputs instead usually produces an out-of-distribution input the network has never seen. - **Dropout is a training-time device only.** At sampling you deliberately evaluate both branches; you do not randomly drop the condition. - **The dropout rate is a trade.** Too low and the unconditional branch is undertrained, making the difference vector noisy; too high and conditional capacity is wasted on examples that carry no condition. - **It generalises past class labels.** Anything that can be embedded - a text encoding, a reference vector, a robot goal pose - can be dropped and nulled the same way, which is why the technique became the default conditioning mechanism across domains. ## Common misreadings Candidates often say the unconditional branch is 'a second model'. It is the same weights with a different conditioning input. Others claim the method needs no extra compute at sampling; it needs double. Others believe the conditioning dropout happens at generation time; it happens only during training. And some describe the mix as an interpolation between the two predictions - it is an interpolation only for `w` between zero and one, and the interesting regime is the extrapolation above one. ## A compact answer One network, conditioning replaced by a learned null token on a small share of training examples, two evaluations per reverse step, and predictions combined as unconditional plus `w` times the conditional-minus-unconditional difference - which is an implicit classifier gradient by Bayes' rule, obtained without ever training a classifier.
- How does classifier guidance differ from the classifier-free construction?Classifier guidance requires a separate classifier trained on noisy inputs at every noise level; its gradient of `log p(c | x_t)` is added to the generator's score estimate. Classifier-free guidance drops that model and recovers the same gradient implicitly, because the difference between the conditional and unconditional noise predictions is proportional to it by Bayes' rule. One model instead of two, at the cost of a second evaluation per step.
- Why is the unconditional branch driven by a learned null embedding rather than by zeroing the conditioning input?Zeros are an input the network has essentially never encountered, so its behaviour there is undefined and the resulting difference vector is unreliable. A learned null embedding is trained on real examples during conditioning dropout, so the unconditional prediction is a proper estimate of the unconditional score and the difference between branches means what the derivation says it means.
- What goes wrong if the conditioning-dropout rate is set far too high?Too large a share of training examples carry no condition, so conditional capacity is underused and the conditional branch weakens - conditioning adherence drops even before any sampling-time mixing. The rate is a balance: enough dropped examples to fit a solid unconditional branch, few enough that the conditional path still sees the overwhelming majority of the data.
saying these in an interview costs you the question
- Says the method needs a separately trained classifier
- Claims the unconditional branch is a second network
- Forgets that sampling costs two evaluations per step
- Thinks conditioning dropout is applied at generation time
- Describes the mix as interpolation only, missing the extrapolation