Why is minimising squared error the same as maximum likelihood under Gaussian noise?
answer
- write down the noise model first
- log of a Gaussian density
- the exponent is the squared residual
- constants cannot move the minimiser
- log-loss is the Bernoulli version
basics
~20 sAssume the errors are independent Gaussian with constant variance. The log-likelihood then equals a constant minus the sum of squared residuals divided by twice the variance, so maximising it and minimising squared error give the same fitted parameters.
solid answer
~40 sWrite the model as `yi = f(xi) + ei` with `ei` independent Gaussian, mean zero, constant variance `sigma^2`. Each observation contributes `-log(sqrt(2*pi*sigma^2)) - (yi - f(xi))^2 / (2*sigma^2)` to the log-likelihood, so summing gives `l = constant - (1/(2*sigma^2)) * sum((yi - f(xi))^2)`. The only term touching the model parameters is the residual sum of squares, entering with a negative sign, so maximising the likelihood is exactly minimising squared error -- and `sigma^2` scales the objective without moving its minimiser. The same logic gives log-loss: for binary labels with predicted probability `pi`, the Bernoulli log-likelihood is `sum(yi*log(pi) + (1-yi)*log(1-pi))`, and its negative is precisely cross-entropy. The practical consequence is that choosing a loss is choosing an error model, so squared error quietly assumes symmetric, constant-variance, light-tailed noise.
go deeper
Know the headline: squared error comes from assuming Gaussian errors and log-loss from assuming Bernoulli labels. Being able to say that a loss encodes an assumption is enough at this stage.
Be able to carry out the derivation -- log the Normal density, sum over observations, and show the residual sum of squares is the only parameter-dependent term, scaled by a constant that cannot move the optimum.
Demonstrate that you use the framing operationally: reading residual asymmetry, changing spread or fat tails as evidence that the assumed error distribution is wrong, and reaching for a different likelihood rather than trimming inconvenient points.
Own the choice of objective across a team's models. Argue when the default loss is defensible, when the error structure calls for a different likelihood, and how that decision is documented so results stay comparable and reviewable.
## The derivation Suppose the data are generated as `yi = f(xi) + ei`, with `ei` independent and Normal with mean 0 and variance `sigma^2` Here `f` is whatever predictor the model produces, with parameters to be fitted. Conditional on the inputs, each `yi` is Normal with mean `f(xi)` and variance `sigma^2`, so its density is `(1 / sqrt(2*pi*sigma^2)) * exp(-(yi - f(xi))^2 / (2*sigma^2))` Taking logs and summing over `n` independent observations: `l = -(n/2)*log(2*pi*sigma^2) - (1/(2*sigma^2)) * sum((yi - f(xi))^2)` The first term does not involve the model parameters at all. The second is the residual sum of squares, multiplied by the negative constant `-1/(2*sigma^2)`. Maximising `l` over the parameters is therefore identical to minimising `sum((yi - f(xi))^2)`. Because `sigma^2` enters only as a positive scale factor on that sum, it does not change which parameters win -- you can fit by least squares without knowing the noise level. If you do want it, profiling `sigma^2` out of the same log-likelihood returns the residual sum of squares divided by `n`, which is the Normal variance MLE applied to residuals. ## Log-loss is the same story with Bernoulli labels Now take binary labels `yi` in {0,1} and a model producing a predicted probability `pi` for each observation. The Bernoulli probability of the observed label is `pi^yi * (1-pi)^(1-yi)`, so the log-likelihood is `l = sum( yi*log(pi) + (1-yi)*log(1-pi) )` The negative of this, usually averaged over observations, is exactly binary cross-entropy, better known as log-loss. So minimising log-loss is not a heuristic borrowed from information theory and separately convenient -- it is maximum likelihood for a Bernoulli response, written with a minus sign because optimisers minimise. This also explains why log-loss punishes confident mistakes so brutally. A predicted probability of 0.01 for an observation whose label is 1 contributes `-log(0.01) ~ 4.6`, and a prediction of 0.001 contributes about 6.9. That is not an arbitrary penalty schedule; it is the log-likelihood of an event the model called nearly impossible. ## The real point: a loss is a distributional assumption The pattern generalises. Common training losses are negative log-likelihoods of some assumed error distribution: - Squared error corresponds to Gaussian errors with constant variance. - Absolute error corresponds to double-exponential (Laplace) errors, which have heavier tails, and its optimum tracks a conditional median rather than a conditional mean. So the question "which loss should I use?" is the question "what do I believe the errors look like?" wearing different clothes. That reframing is what a senior interview is fishing for. ## Where the Gaussian assumption bites in practice Three failure modes are worth being able to name. **Heavy tails and outliers.** The squared-error objective grows quadratically in the residual, so one observation ten times further out contributes a hundred times the pull. Under a genuinely Gaussian model that is correct behaviour -- such a point is nearly impossible, so it should dominate. If the true error distribution has heavy tails, the point is not nearly impossible, and squared error over-reacts to it. A heavier-tailed likelihood, and hence a gentler loss, is the principled fix rather than deleting the point. **Non-constant variance.** The derivation assumed one `sigma^2` shared across observations. When the spread of the errors grows with the level of the response, that assumption is wrong, and unweighted least squares implicitly over-trusts the noisy observations. Allowing an observation-specific variance in the likelihood produces weights inversely proportional to those variances, which is exactly weighted least squares. **Bounded or discrete responses.** Fitting squared error to counts or to 0/1 labels assumes a symmetric, unbounded error distribution around a mean that the model may push outside the feasible range. The Bernoulli or count likelihood is the honest choice, and its negative log gives a loss that stays finite and well-behaved on the right support. ## Diagnosing it Since the loss encodes a noise model, residual behaviour is the check on that model. Look at the residuals for asymmetry, for spread that changes with the fitted value, and for a tail far fatter than a Normal would produce. Each of those is a statement that the likelihood you optimised is not the likelihood the data came from -- and the value of the framing is that it turns a vague complaint about model fit into a specific, testable claim about the error distribution you assumed. ## Saying it in an interview "Under independent Gaussian errors with constant variance the log-likelihood is a constant minus the residual sum of squares over twice the variance, so maximising it is minimising squared error. The same argument makes log-loss the Bernoulli log-likelihood. That means picking a loss is picking a noise model, and I would defend squared error by checking that the residuals actually look symmetric, homoscedastic and light-tailed."
- Does the fit change if the noise variance sigma-squared is unknown?No. In the log-likelihood `sigma^2` multiplies the residual sum of squares by a positive constant and appears otherwise only in a term free of the model parameters, so it cannot move the minimiser. You fit the parameters by least squares regardless, then estimate the variance by profiling: the residual sum of squares divided by `n`.
- Which noise model corresponds to minimising absolute error instead of squared error?A double-exponential, or Laplace, error distribution. Its log-density is a constant minus the absolute residual over a scale parameter, so maximising that likelihood minimises the sum of absolute residuals. Its heavier tails mean far-out observations exert bounded pull, and the optimum tracks a conditional median rather than a conditional mean.
- Why does log-loss penalise a confident wrong prediction so heavily?Because it is the negative log-likelihood of the observed label. Predicting probability 0.01 for an observation that turns out to be a 1 contributes `-log(0.01)`, about 4.6, and 0.001 contributes about 6.9. The penalty is unbounded as the predicted probability approaches zero, because the model declared an event that actually happened to be nearly impossible.
- How would you check whether the Gaussian error assumption behind squared error is reasonable?Examine the residuals. Look for symmetry around zero, spread that stays constant as the fitted value changes, and tails no heavier than a Normal's. Skew, a spread that widens with the fitted value, or a handful of residuals many standard deviations out each say the assumed likelihood differs from the one that generated the data.
saying these in an interview costs you the question
- Treats squared error as assumption-free rather than a noise model
- Says the noise variance must be known before fitting
- Calls log-loss unrelated to any likelihood
- Deletes outliers instead of questioning the assumed error distribution
- Fits squared error to binary labels without noticing the assumption