For an autoencoder's reconstruction loss, when do you use cross-entropy over squared error?
answer
- every loss is a log-likelihood
- what output distribution does it assume
- check the target range first
- Gaussian versus Bernoulli decoder
- sigmoid plus cross-entropy gives p minus x
basics
~20 sUse Bernoulli cross-entropy when every output lies in [0,1] and the decoder emits it through a sigmoid, such as scaled pixels. Use squared error for unbounded real-valued outputs. The choice declares what output distribution you are assuming.
solid answer
~50 sEach reconstruction loss is a negative log-likelihood in disguise, so picking one is picking a decoder output distribution. Squared error corresponds to a Gaussian with fixed variance centred on the decoder output, which is the right assumption for unbounded real values — log-mel spectrogram frames, for instance, are positive and negative and unbounded, so cross-entropy is not even defined on them. Bernoulli cross-entropy corresponds to each output being an independent probability in [0,1], which fits pixels scaled to [0,1] behind a sigmoid. With that pairing the gradient with respect to the pre-activation is simply `p - x`, so a badly wrong saturated output still gets a full-strength gradient, whereas squared error through a sigmoid multiplies by the sigmoid derivative and nearly stalls. One caveat: with fractional targets cross-entropy bottoms out at the target's own entropy rather than zero, so its absolute value means little and is not comparable across losses.
go deeper
Know that the loss must match the range of the data: values squeezed into [0,1] behind a sigmoid pair with cross-entropy, unbounded real values pair with squared error.
Derive both losses as negative log-likelihoods of a Gaussian and a Bernoulli decoder, and explain why the sigmoid-plus-cross-entropy gradient reduces to prediction minus target while squared error stalls in saturation.
Show judgment about scaling and weighting: which channels dominate a squared-error sum, how you handle a record with mixed continuous and binary fields, and why absolute loss values across losses tell you nothing.
Own the framing that the reconstruction loss is a modelling assumption about the decoder's output distribution and about which errors your product can tolerate, and be able to justify a deliberately weighted loss to a team that wanted a default.
## Both losses are log-likelihoods A reconstruction loss is not an arbitrary distance. Writing the decoder as a conditional distribution over outputs given the code, and minimising the negative log-likelihood of the true input under it, produces the familiar losses directly. Assume each output dimension is Gaussian with mean equal to the decoder output and a fixed shared variance. The negative log-likelihood of one dimension is `(x - x_hat)^2 / (2*sigma^2)` plus a constant that does not depend on the parameters. Minimising it is minimising squared error. The fixed variance is why plain squared error weights every dimension equally, and it is the assumption most people make without noticing. Assume instead each output dimension is Bernoulli with probability equal to the decoder output. The negative log-likelihood of one dimension is `-(x*log(x_hat) + (1-x)*log(1-x_hat))` — binary cross-entropy. This requires `x_hat` strictly inside (0,1), which is why it is paired with a sigmoid output, and requires the target `x` to lie in [0,1] for the expression to be meaningful. Both also assume the output dimensions are **conditionally independent given the code**: the total loss is a sum over dimensions with no term coupling neighbouring pixels or adjacent frames. All structure has to arrive through the code; the loss itself never rewards getting the relationship between two outputs right, only each one separately. ## When each one applies **Bernoulli cross-entropy** fits data that is genuinely bounded to [0,1] and is naturally read as an intensity or a probability. Handwritten-digit pixels rescaled from 0-255 to [0,1] are the canonical case. Note that the targets need not be 0 or 1 — the expression is perfectly well behaved at `x = 0.37`, is still minimised at `x_hat = x`, and is a valid training signal. **Squared error** fits unbounded real values: log-mel spectrogram frames, standardised sensor readings, embeddings, residuals. Applying cross-entropy here is not a stylistic mistake but an undefined one — `log(x_hat)` weighted by a target of `-3.1` has no meaning, and clipping the data into [0,1] first destroys exactly the magnitude information you were trying to reconstruct. ## The gradient argument There is a second, practical reason to prefer the matched pairing. Let `z` be the decoder's pre-activation and `p = sigmoid(z)`. With cross-entropy, `dL/dz = p - x`. The sigmoid derivative cancels analytically, so an output that is confidently wrong — `p` near 1 when `x` is 0 — produces a gradient near 1, the largest correction available. With squared error through the same sigmoid, `dL/dz = (p - x) * p * (1 - p)`. The second factor approaches zero exactly when the output is saturated, so the most wrong outputs generate the smallest gradients and training crawls out of saturation slowly. This is the same saturation argument that makes cross-entropy the default for classification, applied to reconstruction. ## Reading the numbers The two losses are not on a common scale and their floors differ. Squared error reaches zero at a perfect reconstruction. Cross-entropy with a fractional target does not: at `x_hat = x = 0.37` the per-dimension loss equals the binary entropy of 0.37, roughly 0.66 nats, and the total loss floor is the sum of those entropies over all dimensions. A model reporting a cross-entropy of 0.7 per pixel may be at the optimum or far from it; you cannot tell from the absolute number. Compare within a loss, across models trained on the same data, never across losses. ## Scaling, weighting and the choice you make implicitly Because squared error assumes a single shared output variance, every dimension contributes in its own raw units. If one sensor channel swings over hundreds and another over hundredths, the loud channel dominates the total error and the code will spend nearly all its capacity there. Standardising each channel before training, or attaching per-dimension weights, is the fix — and either amounts to declaring a per-dimension output variance, which is what the fixed-variance Gaussian assumption was hiding. The same logic covers mixed inputs. A record with continuous and binary fields wants a per-field loss — squared error for the continuous ones, cross-entropy for the binary ones — summed with weights you choose deliberately, because the relative weighting decides which fields the bottleneck prioritises. ## The short version for an interview Say what distribution you are assuming, check that the target range is compatible with it, and mention the gradient behaviour of sigmoid plus cross-entropy. Then add the caveat about comparing loss values across losses. That sequence answers the question and shows you know the losses are modelling choices rather than defaults.
- Can you use Bernoulli cross-entropy when the pixel targets are fractional, like 0.37?Yes. The expression is well defined for any target in [0,1] and is still minimised at a prediction equal to the target, so it trains correctly. The catch is the floor: at the optimum the loss equals the target's binary entropy rather than zero, so the absolute loss value carries no information about how good the reconstruction is.
- Why does a sigmoid output trained with squared error learn slowly?The gradient with respect to the pre-activation carries a factor of the sigmoid derivative, which goes to zero as the output saturates. So an output that is confidently wrong receives almost no gradient. Pairing the sigmoid with cross-entropy cancels that factor analytically and leaves the gradient equal to prediction minus target.
- What goes wrong with squared error when input channels have wildly different scales?Squared error assumes one shared output variance, so each dimension contributes in its raw units and the highest-amplitude channels dominate the total. The bottleneck then spends its capacity reconstructing the loud channels and ignores the quiet ones. Standardise per channel before training, or attach explicit per-dimension weights, which is the same thing as declaring per-dimension variances.
saying these in an interview costs you the question
- Says cross-entropy is only for classification, never reconstruction
- Applies cross-entropy to unbounded or negative-valued targets
- Compares loss values across the two losses as if commensurable
- Cannot connect a reconstruction loss to an assumed output distribution
- Treats the loss choice as mere numerical convenience