Why does a VAE's encoder emit log-variance rather than variance or standard deviation?
answer
- a linear head can output any real number
- variance carries a hard constraint
- the useful range spans many orders of magnitude
- the KL formula already contains a log
basics
~20 sLog-variance is unconstrained, so any real number a linear output emits is legal, and exponentiating it always yields a positive variance. It also keeps tiny variances representable and drops straight into the Gaussian KL formula, which already contains a log.
solid answer
~40 sThe encoder's heads are ordinary linear outputs, so they can emit any real number, positive or negative — but a variance must be strictly positive. Predicting the log of the variance makes the constraint free: `sigma^2 = exp(logvar)` is positive for every possible output, with no clipping, absolute value or softplus in the way. It is also numerically kinder than predicting the standard deviation: a near-deterministic posterior shows up as a large negative log-variance rather than as a value underflowing to zero, and the closed-form KL term `0.5 * sum(mu^2 + exp(logvar) - 1 - logvar)` consumes the log-variance directly, so no `log` of a possibly-zero number is ever taken. In practice the emitted log-variance is clamped to a sane range, because `exp` of a large positive value overflows.
go deeper
Recall that the encoder emits two vectors per input, a mean and a spread, and that the spread is emitted in log form so that exponentiating it can never produce a negative or zero variance.
Explain all three reasons: unconstrained output, a log scale that keeps very small variances representable without underflow, and a closed-form KL that already consumes a log-variance. Be able to recover sigma as exp(half the log-variance).
Show you have debugged this: clamping the emitted log-variance to stop exp from overflowing, and reading heads pinned at the bounds as deterministic or unused dimensions rather than as a healthy fit.
Frame it as a general design habit — choose the parameterisation in which the constraint is structural and the optimiser takes additive steps, rather than bolting a positivity map and a clip onto an ill-scaled output.
## The constraint problem A VAE's encoder produces the parameters of a per-input Gaussian over latent codes: a mean vector and a spread. The mean is unconstrained — any real number is a legal mean, so a plain linear output head works. The spread is not: a variance must be strictly greater than zero, and a standard deviation likewise. A linear head has no such restriction; it will happily emit `-3.1`. So one of three things has to happen. Either the network's raw output is mapped through something that guarantees positivity, or the parameterisation is chosen so that positivity is automatic, or training has to survive occasional illegal values. The standard choice is the second: the head predicts `logvar = log(sigma^2)`, a quantity for which every real number is meaningful, and the variance is recovered as `sigma^2 = exp(logvar)`, with the standard deviation as `sigma = exp(0.5 * logvar)`. ## Why not predict the standard deviation through a positivity map? You can — passing a raw output through a softplus or an absolute value produces a positive number too — but each option has a cost the log parameterisation avoids. - **Absolute value or squaring** puts a kink or a stationary point at zero. Gradients behave badly exactly where the model is trying to express near-certainty. - **Softplus** is smooth, but its gradient vanishes as the output goes very negative, so the model gets slow to drive a dimension toward tiny variance. - **Clipping a raw output at a small positive floor** kills the gradient outright whenever the clip is active. More importantly, all of these live on a *linear* scale, where the interesting range spans orders of magnitude. A posterior variance of `1e-6` and one of `1e-3` are very different states of the model, and on a linear scale they are numerically adjacent to each other and to zero. On the log scale they are `-13.8` and `-6.9` — comfortably separated, comfortably representable, and reachable by additive steps of the kind gradient descent naturally takes. Multiplicative structure in the variance becomes additive structure in the log, which is what an optimiser is good at. ## The KL term wants the log anyway The regularisation term for a diagonal-Gaussian posterior against a standard-normal prior is, per latent dimension: ``` 0.5 * ( mu^2 + sigma^2 - 1 - log sigma^2 ) ``` Note the `log sigma^2` sitting there. If the head emits the log-variance directly, the whole expression is `0.5 * (mu^2 + exp(logvar) - 1 - logvar)` and no logarithm is ever computed at runtime — which matters, because `log` of a variance that has been driven to exactly zero is negative infinity, and a single such value poisons the whole batch's loss. Emitting the log makes that failure mode unreachable by construction: `exp` never returns exactly zero for a finite input. ## The practical caveat The parameterisation is not free of numerical hazards, it just moves them. `exp(logvar)` overflows if the head emits a large positive number, which can happen during an unstable early epoch. The usual guard is to clamp the emitted log-variance to a range — something like `[-7, 7]` covers everything a healthy model needs, since the upper end is a variance in the thousands and the lower end a standard deviation around `0.03`. Clamping also gives a useful diagnostic: a head pinned at the lower bound is a posterior that has gone effectively deterministic on that dimension, and a head pinned at the upper bound usually means the dimension is unused and has been parked on the prior. ## What an interviewer is checking The question is a quick test of whether you have actually built one of these rather than read about it. The answer has three parts, and a candidate who gives only the first is giving half of it: positivity comes for free, the log scale spans the range the model needs without underflow, and the closed-form KL already speaks in log-variance so nothing extra has to be computed. Mentioning the clamp shows the third layer — that you have watched one of these blow up.
- What goes wrong if the encoder predicts the variance directly?You need a positivity map, and each one costs something: squaring or an absolute value puts a kink at zero, softplus starves the gradient when the model wants a tiny variance, and clipping at a floor kills it entirely. Worse, a variance that reaches exactly zero makes the `log sigma^2` inside the KL term infinite.
- Why do practitioners clamp the emitted log-variance to a range?Because `exp` of a large positive output overflows during unstable training, and an extremely negative output means a posterior that has gone deterministic on that dimension. A clamp of roughly plus or minus seven covers every variance a healthy model needs, and heads pinned at either bound are a useful diagnostic.
- How do you get the standard deviation back from the emitted log-variance?`sigma = exp(0.5 * logvar)`, since the head emits the log of the variance rather than the log of the standard deviation. Getting the factor of one half wrong is a common bug: it silently squares or square-roots the noise scale, so the KL and the sampled codes disagree about the posterior width.
saying these in an interview costs you the question
- Says it is only a notational convenience
- Thinks the encoder emits a variance and logs it later
- Forgets the factor of one half converting to a standard deviation
- Claims a full covariance matrix must be emitted