skip to content

In diffusion image generation, what does raising the guidance scale trade away?

level: middleimportance: must knowfreq 62%

answer

  1. one dial, two competing goods
  2. prompt adherence versus natural variety
  3. runs the denoiser twice per step
  4. past a point, colour and geometry break
  5. model-specific band, always sweep

basics

~20 s

Guidance scale trades diversity and naturalness for prompt adherence. Low values give loose, varied, plausible images that may ignore parts of the brief; high values track the prompt harder but flatten variety and, past a point, produce over-saturated, contorted output.

solid answer

~50 s

Classifier-free guidance runs the denoiser twice per step — once conditioned on the prompt, once unconditioned — and extrapolates away from the unconditioned prediction. The guidance scale is how far it extrapolates. At 1 there is effectively no guidance and you get whatever the model finds likely; in the mid range the sample is pulled toward the prompt while staying on the natural-image manifold; push it high and you leave that manifold, which shows up as blown contrast, over-saturated colour, waxy texture and distorted geometry. It also collapses diversity: several seeds start producing near-identical compositions. Practically you sweep it. For a furniture catalogue prompt, a scale of 2 gives loose, plausible rooms that drift off-brief, around 7 tracks the brief while looking photographic, and 15 gives lurid, contorted furniture. The usable band is model-specific, and some distilled few-step models bake guidance in and ignore the knob entirely.

code

python · 7 lines
python
# Classifier-free guidance, one sampling step
eps_uncond = denoiser(x, t, cond=empty_prompt)
eps_cond   = denoiser(x, t, cond=prompt)

# scale == 1.0 -> plain conditional; larger -> extrapolate past it
eps = eps_uncond + scale * (eps_cond - eps_uncond)
x = step(x, eps, t)

go deeper

for a junior

Know that guidance scale controls how strictly the image follows the prompt, and that turning it up too far makes pictures look garish and warped rather than simply better.

for a middle

Explain the two-prediction extrapolation behind classifier-free guidance, name the symptoms at low, tuned and excessive values, and separate guidance failures from step-count failures.

for a senior

Demonstrate the operational habit: sweep per model, record guidance with seed and version as part of a reproducible config, and recognise a garish brand asset as an over-guided pipeline rather than a prompt problem.

for a principal

Own the cost angle — guidance doubling per-step compute, and guidance-distilled variants changing that economics — and set the policy that adherence problems are solved with better conditioning, not by cranking a knob that degrades realism.

## What the dial actually is Guidance scale — often called CFG scale, for classifier-free guidance — is the strength of the pull the text prompt exerts on each denoising step. Mechanically, the model produces two predictions at every step: one with the prompt as conditioning, and one with an empty or negative conditioning. The sampler then takes the unconditional prediction and moves *past* the conditional one, by a factor equal to the guidance scale, along the difference between them. A scale of 1 means "just use the conditional prediction". A scale of 7 means "go seven times further in the direction the prompt pushed". The cost of this is immediate and worth stating in an interview: classical classifier-free guidance roughly doubles the compute per step, because the denoiser runs twice. Some deployments amortise that by batching the two passes together. ## The fidelity-diversity dial The extrapolation is why the tradeoff exists. The unconditional prediction points toward "any plausible image"; the conditional one points toward "a plausible image matching this text". Extrapolating along their difference amplifies whatever is *distinctive about the prompt* relative to generic imagery. A little amplification sharpens adherence. A lot amplifies everything the prompt implies until the sample leaves the region of parameter space that corresponds to real photographs. Concretely, sweeping a single catalogue prompt — the same armchair rendered in a set of room styles — shows the whole curve: - **Around 2.** Rooms look natural and varied but the brief slips: the wrong wood tone, a different arm profile, the requested season absent. Seeds differ wildly from one another. - **Around 7.** The brief is honoured, the geometry is photographic, and there is still meaningful variety between seeds. This is the band most models are tuned for. - **Around 15.** Colours go neon, contrast blows out, edges get a hard airbrushed look, and object geometry contorts — an armchair grows an extra leg or the perspective of the room bends. Multiple seeds converge on nearly the same over-cooked composition. The last symptom is the one candidates forget. High guidance does not merely add artifacts; it *reduces the effective variety of the model*, which is a real problem when the whole point of the batch is to produce six distinct room styles. ## Interactions worth knowing **With steps.** Guidance and step count are independent knobs and each fixes a different failure. Blurry, undercooked output is usually a step-count problem. Off-brief output is a guidance (or prompt) problem. Adding steps at guidance 15 does not remove the saturation. **With the negative prompt.** The unconditional branch can be replaced by a *negative* prompt, which turns the same extrapolation into "move away from this description". The guidance scale then controls how hard you flee the negative as well as how hard you chase the positive, so a strong negative prompt at high guidance compounds the distortion. **With the architecture.** Model families are tuned for different bands. Older step-heavy models were commonly run in the 5-9 range; several current flow-matching models are tuned for noticeably lower values, and guidance-distilled few-step variants have the guidance behaviour trained into the weights so the runtime knob does nothing or is exposed as a fixed embedded value. Never carry a number across model families — sweep it once per model and record it alongside the prompt. ## How to use it in production Treat guidance as part of the reproducible generation config, next to model version, seed, step count and scheduler. When a brand-consistency pipeline starts producing garish assets, the first thing to check is whether someone raised guidance to force adherence to a longer prompt. The better fix for adherence is almost always a clearer prompt or an explicit structural conditioning signal, not more guidance — guidance is a blunt amplifier and it degrades the very realism the catalogue depends on. ## Diagnostic summary If output ignores the brief: raise guidance a little, or rewrite the prompt to front-load the requirements. If output is saturated, waxy or anatomically wrong across seeds: lower guidance. If output is soft and undercooked: add steps. If every seed looks the same: lower guidance to recover diversity.

  • Why does high guidance reduce variety between seeds, not just add artifacts?
    Guidance amplifies the direction that distinguishes the prompt from generic imagery, so every trajectory is pushed hard toward the same narrow mode of the conditional distribution. Different starting noise then converges on similar compositions. That matters when the batch exists to produce genuinely different options — you are paying for six seeds and receiving one idea rendered six times.
  • A prompt keeps dropping one requirement. Is raising guidance the right fix?
    Usually not. Guidance amplifies the conditioning the text encoder already produced; if the requirement was weakly encoded or buried at the end of a long prompt, amplification distorts the rest of the image before it rescues that detail. Rewrite the prompt to state the requirement early and concretely, or impose it structurally through a conditioning image, and keep guidance in its tuned band.
  • What does guidance cost you at inference time?
    Classical classifier-free guidance evaluates the denoiser twice per step — conditional and unconditional — so it roughly doubles compute per image compared with unguided sampling. That is a real line item at scale, and it is one reason guidance-distilled variants, which fold the behaviour into the weights and need only one pass, are attractive for high-volume or interactive workloads.

It behaves like the contrast slider on a photo: a little makes the subject read clearly, a lot makes it unmistakable but plainly artificial, and everything ends up looking the same.

saying these in an interview costs you the question

  • Treats guidance scale as an overall quality slider
  • Says higher guidance always improves prompt adherence
  • Confuses guidance scale with the number of sampling steps
  • Carries one model's tuned value across to another family
  • Ignores that high guidance collapses diversity across seeds

context