Squared-error loss is maximum likelihood under what assumption about the target?
answer
- A loss encodes a noise model
- Take the negative log of a density
- Constants cannot move an argmin
- Symmetric, same spread everywhere, thin tails
- The minimiser is the conditional mean
basics
~20 sThat the target is Gaussian around the network's output with a constant, input-independent variance. Minimising squared error is exactly maximising that Gaussian likelihood, which is why the trained output estimates the conditional mean of the target.
solid answer
~50 sSquared error is the negative log-likelihood of a Gaussian with mean `mu(x)` — the network's output — and a fixed variance. Writing it out, `-log p(y|x) = (y - mu)^2 / (2 * sigma^2) + 0.5 * log(2 * pi * sigma^2)`; with `sigma` held constant, the second term and the `1/(2 * sigma^2)` factor are constants in the parameters, so minimising squared error and maximising that likelihood have the identical minimiser. Three assumptions ride along: the noise is symmetric, its scale is the same for every input (homoscedastic), and its tails are thin, so a rare extreme target is treated as nearly impossible and gets an enormous penalty. The practical consequence is that a squared-error network learns `E[y|x]`, the conditional mean. On a right-skewed target like delivery time, the mean sits above the typical case, so the predictions are pulled up by the tail — that is the loss doing exactly what it was told, not a bug.
go deeper
Remember that squared error is not a neutral default: it assumes errors are symmetric bell-shaped noise of roughly equal size everywhere, and that the model's output is a prediction of the average target.
Be able to write the Gaussian negative log-likelihood, show that the variance drops out as a scale and an additive constant, and name the conditional mean as the minimiser.
Show which assumption a real dataset breaks and what you did about it: skew that makes the mean the wrong answer for the user, or noise scale that varies wildly across the input space.
Frame loss selection as an explicit modelling decision the team can defend. Decide which statistic the product actually needs from a prediction, and make that choice visible in the objective rather than patched in post-processing.
## Every loss is a distributional claim Choosing a loss is choosing a probability model, whether or not you say so. Squared error corresponds to assuming that, given the input `x`, the target is drawn as ``` y | x ~ Normal(mu(x), sigma^2) ``` where `mu(x)` is what the network outputs and `sigma` is a fixed constant that does not depend on `x` and is not learned. ## The derivation The Gaussian density is `p(y|x) = (1 / sqrt(2 * pi * sigma^2)) * exp(-(y - mu)^2 / (2 * sigma^2))`. Maximum likelihood maximises the log of that over the training set, which is the same as minimising the negative log-likelihood: ``` -log p(y|x) = (y - mu)^2 / (2 * sigma^2) + 0.5 * log(2 * pi * sigma^2) ``` Only the first term contains `mu`, and it is the squared residual multiplied by the positive constant `1/(2 * sigma^2)`. The second term is a constant. Multiplying an objective by a positive constant and adding a constant cannot move its minimiser, so ``` argmin_theta sum_i (y_i - mu(x_i))^2 == argmax_theta sum_i log p(y_i | x_i) ``` Setting `sigma = 1` is therefore a harmless convention for the mean, not an extra restriction on it. (It does matter if you ever want to compare the loss value to a likelihood, or if the variance itself is a quantity you want to learn.) ## What the assumption buys and what it costs The assumption has three moving parts, and each has a visible consequence. **Symmetry.** A Gaussian is symmetric about its mean, so over-predicting by 5 and under-predicting by 5 cost the same. If your problem does not treat those alike, squared error is silently the wrong objective. **Constant variance (homoscedasticity).** Every example's residual is judged on the same scale. When the true noise scale varies across the input space — quiet, predictable regions and loud, unpredictable ones — the loss spends most of its budget on the loud regions simply because their residuals are naturally larger, even when the model is already doing as well as anything could there. Nothing in the objective can express "this region is inherently noisy", because `sigma` is constant by construction. **Thin tails.** Gaussian density decays like `exp(-r^2)`, so a residual ten standard deviations out is assigned an astronomically small probability. Under maximum likelihood that translates into an astronomically large penalty and a correspondingly large gradient. This is the likelihood-side explanation of the same phenomenon you see mechanically as an outlier dominating an update: you told the model such a value was essentially impossible, so it will contort itself to explain it. ## Which statistic you end up predicting The cleanest way to see what a loss does is to ask which constant `c` minimises the expected loss against a random target `y`: - `E[(y - c)^2]` is minimised at `c = E[y]`, the mean; - `E[|y - c|]` is minimised at `c = median(y)`; - Huber's expectation is minimised somewhere between the two, closer to the median as its crossover shrinks. Because a network computes a separate output per input, it targets these conditionally: a squared-error head is trained to output `E[y|x]`, the conditional mean, and an absolute-error head the conditional median. For a symmetric target the two coincide and the choice is mostly about optimisation. For a right-skewed target — delivery time, where most deliveries are quick and a few are very slow — they do not: the conditional mean is pulled up by the slow tail and sits above the typical delivery. Neither answer is wrong; they answer different questions. If a stakeholder asks "how long will this delivery take?" they usually mean the typical case, and a mean-seeking loss will not give it to them. ## Using the framing in practice The likelihood view is not academic bookkeeping. It tells you what to change when squared error misbehaves: - If the *symmetry* assumption is wrong for the decision being made, the objective needs to reflect that asymmetry rather than being patched afterwards. - If the *constant-variance* assumption is wrong, the natural repair is to let the model express a per-input spread instead of pretending one number fits everywhere. - If the *thin-tail* assumption is wrong, a heavier-tailed noise model — of which Huber is the practical, gradient-capped stand-in — is the principled response, and its cost is that you are no longer estimating the mean. In every case the useful question in an interview is not "which loss is best" but "what did this loss assume, and which of those assumptions does my data violate".
- Which noise assumption gives absolute error instead, and which statistic does it target?A Laplace distribution: its density decays like `exp(-|y - mu| / b)`, so the negative log-likelihood is the absolute residual up to constants. Its expected loss is minimised at the median, so an absolute-error head estimates the conditional median. On a right-skewed target the median sits below the mean, which is exactly the gap you see between an absolute-error and a squared-error model on the same data.
- Why is it harmless to fix the assumed variance at one when training the mean?Because the variance enters the objective only as a positive multiplicative factor `1/(2 * sigma^2)` and an additive constant. Neither can change where the minimum over the mean parameters lies. It does change the numerical scale of the loss, which interacts with the learning rate, and it means the reported loss is not comparable to a true log-likelihood.
- How does the thin-tail assumption connect to squared error's fragility on corrupted targets?Gaussian density falls off like `exp(-r^2)`, so a far-out residual is declared nearly impossible. Maximum likelihood then assigns it an enormous penalty and an unbounded gradient, and the fit twists to accommodate it. Assuming a heavier-tailed noise model produces a loss whose gradient saturates, which is the principled reading of why robust losses behave better on dirty targets.
saying these in an interview costs you the question
- Says squared error makes no distributional assumption at all
- Claims it assumes the inputs are normally distributed
- Thinks a squared-error head predicts the most likely single target value
- Believes the fixed variance changes which mean is optimal
- Cannot say which statistic squared error's minimiser targets