Why did diffusion models move from a linear noise schedule to a cosine one?
answer
- how fast does the signal die?
- the last third does nothing
- cumulative term near 0.01 at two-thirds
- gentler middle, slow at both ends
- low resolution needs the gentler curve
basics
~20 sA linear beta schedule destroys nearly all signal two-thirds of the way through the chain, so the final third of the reverse steps teaches the model almost nothing. A cosine schedule lets signal-to-noise fall gradually to the end.
solid answer
~50 sThe schedule decides how much of the image survives at each step, summarised by the cumulative product `alpha_bar_t` in `x_t = sqrt(alpha_bar_t) * x_0 + sqrt(1 - alpha_bar_t) * eps`. Under the original linear beta ramp over 1000 steps, `alpha_bar` is already about 0.01 two-thirds of the way through: the image is gone and every remaining step is noise-to-noise, wasting both training samples and sampling time. The cosine schedule sets `alpha_bar_t` proportional to `cos^2` of a rescaled time, so it drops slowly at the start, roughly linearly in the middle, and slowly again at the end - at that same 67% mark it still sits near 0.24. Useful learning signal is spread across the whole chain instead of crammed into the first third. It matters most at low resolution: at 64x64 the linear ramp over-noises, while pixel redundancy makes the same schedule effectively gentler at higher resolutions.
code
python · 21 linesimport math
T = 1000
beta = [1e-4 + (0.02 - 1e-4) * i / (T - 1) for i in range(T)]
lin, acc = [], 1.0
for b in beta: # linear schedule
acc *= (1.0 - b)
lin.append(acc)
s = 0.008 # cosine schedule
f = lambda t: math.cos(((t / T) + s) / (1 + s) * math.pi / 2) ** 2
cos = [f(t + 1) / f(0) for t in range(T)]
for frac in (0.25, 0.50, 0.67, 0.90):
i = int(frac * T) - 1
print(f"at {frac:.2f} of the chain: alpha_bar linear={lin[i]:.4f} cosine={cos[i]:.4f}")
# at 0.25 of the chain: alpha_bar linear=0.5241 cosine=0.8470
# at 0.50 of the chain: alpha_bar linear=0.0786 cosine=0.4938
# at 0.67 of the chain: alpha_bar linear=0.0105 cosine=0.2420
# at 0.90 of the chain: alpha_bar linear=0.0003 cosine=0.0241go deeper
Be ready to say what a noise schedule is: the amount of noise added at each step of the forward process, and therefore how much of the image survives at each step. Knowing that linear and cosine are two common choices is the floor.
An interviewer at this level expects you to explain the cumulative surviving-signal term, show that a linear ramp drives it near zero well before the end of the chain, and say concretely what the cosine curve does differently at the start, middle and end.
Demonstrate that you have diagnosed this in practice: how you would detect wasted high-noise steps, why the fix is a reweighting of where compute lands rather than a bigger model, and how you would validate a schedule change without a full retraining sweep.
Own the framing that a schedule is a compute-allocation decision tied to the data distribution and resolution. Be ready to argue when a schedule change is worth a retrain versus when a loss reweighting or a better sampler gets the same result cheaper.
## What a noise schedule actually is A diffusion model is trained by progressively corrupting a clean sample `x_0` into pure Gaussian noise over `T` steps. Each step adds a little noise with variance `beta_t`, and the standard closed form for the corrupted sample at step `t` is ``` x_t = sqrt(alpha_bar_t) * x_0 + sqrt(1 - alpha_bar_t) * eps, alpha_bar_t = prod_{i<=t} (1 - beta_i) ``` The *schedule* is the sequence `beta_1 ... beta_T`, and the quantity that actually matters is the cumulative product `alpha_bar_t`. It is the fraction of the clean signal's amplitude still present at step `t`. The signal-to-noise ratio at that step is `alpha_bar_t / (1 - alpha_bar_t)`. A schedule is therefore a decision about **how the SNR falls from very high to essentially zero** across the chain. ## Why the linear schedule wastes the chain The original formulation ramped `beta_t` linearly from about 1e-4 to about 0.02 over 1000 steps. Multiply those out and the cumulative signal collapses fast: roughly 0.52 at a quarter of the chain, 0.08 at the halfway point, 0.01 at two-thirds, and 0.0003 at 90%. In amplitude terms the image is unrecognisable well before the halfway mark. That has two consequences. 1. **Wasted training capacity.** Training samples a timestep uniformly and asks the network to denoise there. If most of the timestep range corresponds to inputs that are indistinguishable from pure noise, most sampled training steps present a target the network can nearly satisfy without looking at the image content. The model spends its optimisation budget on a regime that carries almost no information about the data. 2. **Wasted sampling steps.** During generation, the early (highest-noise) reverse steps barely change the sample in any way that survives to the output. Empirically, a large chunk of the reverse chain under a linear schedule can be skipped with very little change in sample quality - which is a direct measurement that those steps were not doing work. ## What the cosine schedule changes The cosine schedule defines the cumulative term directly rather than defining `beta` and multiplying out: ``` alpha_bar_t = f(t) / f(0), f(t) = cos^2( ((t/T + s) / (1 + s)) * pi/2 ) ``` with a small offset `s` (a value near 0.008 is standard) that stops the very first step from being infinitesimally small. Because `cos^2` is flat near its endpoints and steep in the middle, `alpha_bar_t` decays slowly at the start, roughly linearly through the middle, and slowly again as it approaches zero. At two-thirds of the chain it is still around 0.24 rather than 0.01 - there is real image content left for the model to work with, so the late steps are genuine learning problems. The per-step `beta` implied by this curve is derived as `beta_t = 1 - alpha_bar_t / alpha_bar_(t-1)`, and it is clipped (a cap near 0.999 is common) so the final steps stay numerically well behaved. ## Why resolution is part of the answer A schedule is not universally good or bad; it is good or bad *for a data distribution*. Adding independent noise of a fixed variance to a 64x64 image destroys much more semantic content than adding the same noise to a 256x256 image, because neighbouring pixels in the larger image are highly redundant - averaging or downsampling recovers structure that the noise has not independently erased. So the same schedule is effectively **gentler at higher resolution**. That cuts both ways. The linear ramp was tuned where it looked reasonable and is clearly too aggressive at 64x64 and below, which is where the cosine schedule was introduced and where it helps most. Going the other way, when a model is scaled up to a much larger resolution, practitioners deliberately shift the schedule to be *noisier*, because otherwise even the highest-noise steps still leak the global layout and the model never learns to invent composition from scratch. ## How to talk about it in an interview The crisp framing is: the schedule is a budget allocation. You have `T` steps and a fixed amount of learning and sampling compute; the schedule decides how many of those steps land in each SNR band. A linear ramp spends most of its steps in a band where nothing can be learned. A cosine ramp spreads them out. Everything else - what the network predicts, whether sampling is stochastic, how many steps you take at inference - is a separate axis. A common misconception worth pre-empting: changing the schedule does not change the architecture and does not change what the denoiser is asked to predict. It changes only the noise levels the network sees during training and the noise levels the sampler walks back down.
- How does image resolution change which noise schedule you want?Neighbouring pixels in a high-resolution image are redundant, so a fixed noise variance destroys less semantic content there than at 64x64. The same schedule is effectively gentler as resolution rises. That is why the linear ramp is tolerable at 256x256 but over-noises small images, and why scaling a model up to much larger images usually means shifting the schedule to be noisier so the top of the chain really does erase layout.
- If the late steps are wasted, why not simply truncate the chain instead of changing the schedule?Truncating means your sampler starts from a distribution that is not standard Gaussian noise, so you need a learned or approximated prior at the truncation point - extra machinery for the same benefit. Reshaping the schedule keeps the endpoint as plain noise and redistributes where the model spends capacity. Truncation also does not recover the wasted training signal, since training samples timesteps across the whole range.
- How does the schedule affect the effective weighting of the training loss across timesteps?Training draws a timestep uniformly, so the schedule decides how many draws land in each signal-to-noise band. A schedule that collapses the signal early puts most of its draws in a regime where the target is nearly trivial, which is an implicit loss weighting that starves the informative mid-SNR region. Reshaping the schedule is one way to reweight; explicitly reweighting the per-timestep loss is another.
It is like a fade-to-black in a film. A linear schedule goes fully black a third of the way in and then holds on black; the cosine schedule keeps fading gradually so every frame of the transition still shows you something.
saying these in an interview costs you the question
- Says the cosine schedule is just a smoother-looking curve with no measurable effect
- Thinks changing the schedule changes the denoiser architecture
- Claims the schedule only affects sampling, not training
- Assumes one schedule is optimal at every image resolution
- Confuses the per-step beta with the cumulative surviving-signal term