How does a text-to-image diffusion model turn random noise into an image?
answer
- start from static, not a blank page
- many small passes, one network
- the prompt steers every step
- the network estimates the noise
- seed decides which sample you get
basics
~20 sGeneration starts from pure random noise. The model repeatedly predicts how much noise is present and removes a little of it, conditioned on the text prompt at every step. After a fixed number of steps the noise has been shaped into an image.
solid answer
~40 sDiffusion models are trained by taking real images, adding noise at many different strengths, and learning to predict the noise that was added. At generation time you run that in reverse: begin with a tensor of pure random noise, and at each sampling step ask the network what noise it sees, subtract a fraction of it according to a schedule, and repeat. The text prompt is encoded once and fed into every step, so it steers the whole trajectory rather than just the beginning. Nothing is retrieved or pasted from training data — the picture emerges because each step nudges the sample toward the region of image space that the prompt describes. The number of steps is a knob: more steps means a smaller, more careful nudge each time, at proportionally more compute.
code
python · 9 lines# Shape of the sampling loop, framework-agnostic pseudocode
x = random_noise(shape=(1, C, H, W), seed=1234)
cond = text_encoder("a walnut armchair in a sunlit reading room")
for t in schedule(num_steps=30): # high noise -> low noise
noise_pred = denoiser(x, t, cond) # what junk is in x at level t?
x = step(x, noise_pred, t) # remove a fraction of it
image = decode(x)go deeper
Be able to say plainly that generation starts from random noise and that the model repeatedly removes predicted noise while the prompt guides every step. Knowing that nothing is copied from training images is the key point.
Explain the training objective (predict the added noise at a random noise level) and how sampling inverts it, plus what the scheduler and seed control. Be ready to describe why composition settles early and detail late.
Show you can reason about the loop operationally: step count against latency budget, seed pinning for reproducible approved assets, and diagnosing whether a bad output is a step-count problem, a conditioning problem or a prompt problem.
Own the framing that step count is a direct compute-per-image lever and that few-step distilled variants change the cost curve of a whole product surface. Be ready to argue when an interactive experience should trade quality for a sub-second loop.
## The starting point is noise, not a canvas A diffusion image generator does not paint. It does not draw a subject and then a background, and it does not build the image left to right or top to bottom the way an autoregressive text model emits tokens. It begins with a tensor of random values — visually, television static — and progressively refines *the entire image at once* over a fixed number of passes. ## What the model was trained to do Training is the easy half to understand. Take a real image, pick a random noise level, add that much Gaussian noise to it, and hand the noisy image plus the noise level to a network. The network's job is to predict the noise that was added. Because the clean image and the noise are both known during training, this is ordinary supervised learning with a mean-squared-error loss. Do this over hundreds of millions of image-caption pairs at every noise level, and you get a network that, for any noisy image, can say "here is roughly the junk in this, and here is roughly the picture underneath". Some model families parameterise the target slightly differently — predicting the clean image, or a velocity term, rather than the raw noise — but they are all equivalent views of the same denoising signal. ## Sampling: running it backwards Generation inverts the process. Start from pure noise, which is the highest possible noise level. Ask the network what the noise is. Remove a *fraction* of it, taking the sample from "100% noise" to something like "96% noise", and repeat. Each step lands you at a slightly lower noise level until, at the last step, you have a clean image. The reason it is a loop and not a single forward pass is that the network's estimate from pure noise is extremely vague — it can only guess at broad structure. As some noise is removed, the partially formed image gives the network far more to condition on, and its next estimate is sharper. Composition and layout settle in the early steps; texture, edges and fine detail arrive late. This is why an aborted run at step 5 of 30 looks like a coloured blur with the right general arrangement. ## Where the prompt enters The text prompt is encoded once by a text encoder into a sequence of vectors. Those vectors are supplied to the denoiser at *every* step, usually through cross-attention or joint attention inside the backbone. That is the crucial point for interviews: conditioning is not a one-time seed, it is a continuous pull applied to every nudge. It is also why prompt adherence can be dialled up or down independently of step count — that is the guidance scale, a separate control. ## The scheduler A scheduler (or sampler) decides *how many* steps there are and how much noise comes off at each one. Cutting steps makes each removal larger and coarser; below the model's usable range you get mushy or malformed output. Raising steps far above it buys almost nothing and costs linear time. Modern models are tuned for a specific band, and few-step distilled variants are trained specifically to work at a handful of steps. ## Why two runs of the same prompt differ The starting noise is drawn from a random seed. Fix the seed, the prompt, the model, the step count and the guidance value, and generation is essentially reproducible; change the seed and you get a different sample of the same described distribution. This is the standard mechanism for "give me four options" and for reproducing a specific result later — you keep the seed. ## What this rules out Because the network only ever outputs a noise estimate, there is no database lookup and no collage step. The model's knowledge lives entirely in its weights. And because the whole frame is refined jointly, there is no notion of "finishing" one region before another — which is exactly why localized edits need a separate mechanism such as masking, rather than just asking for a smaller change.
- What actually changes between step 3 and step 25 of a 30-step run?Early steps decide low-frequency content: composition, colour blocks, the rough position and pose of objects. Later steps resolve high-frequency content: edges, texture, small text and fine ornament. That ordering is why a prompt change mostly reshuffles layout while a step-count change mostly affects crispness, and why aborting early gives a correctly arranged blur rather than a half-drawn picture.
- If you cut sampling steps from 30 to 6 on a model tuned for 30, what do you see?Each step has to remove far more noise than the model was calibrated for, so errors compound: soft, undercooked textures, mangled small details like hands and lettering, and sometimes a washed-out or over-smooth look. The fix is not more guidance but either restoring the step count or switching to a variant explicitly distilled for few-step sampling.
- Why is the same prompt not reproducible by default?The initial tensor is drawn from a random seed, and each seed picks out a different sample from the distribution the prompt describes. Pin the seed along with the model version, step count and guidance value and runs become repeatable, which is how teams reproduce an approved asset or compare two prompt variants fairly.
It is closer to a sculptor removing marble than a painter adding strokes: the whole block is worked over and over, with the shape emerging everywhere at once rather than being drawn region by region.
saying these in an interview costs you the question
- Says the model retrieves and stitches together training images
- Thinks the image is drawn region by region or left to right
- Claims a single forward pass produces the final picture
- Believes the prompt only influences the first step
- Assumes identical prompts must produce identical images