skip to content

What are the maximum likelihood estimates of the mean and variance of a Normal sample?

level: middleimportance: should knowfreq 52%

answer

  1. two unknowns, two score equations
  2. the variance cancels in the mean equation
  3. mean equation collapses to sum of deviations
  4. substitute the fitted mean back in
  5. maximising returns the divisor n

basics

~20 s

The mean estimate is the sample average xbar. The variance estimate is the average squared deviation from xbar, that is sum of (xi - xbar)^2 divided by n. Maximising the likelihood returns the divisor n, not n-1.

solid answer

~40 s

For `x1, ..., xn` drawn independently from a Normal distribution, the log-likelihood is `l(mu, sigma^2) = -(n/2)*log(2*pi) - (n/2)*log(sigma^2) - (1/(2*sigma^2)) * sum((xi - mu)^2)`. Differentiating with respect to `mu` gives `(1/sigma^2) * sum(xi - mu)`, and setting it to zero yields `mu-hat = xbar`, the sample mean -- note `sigma^2` drops out, so the mean estimate does not depend on the variance. Differentiating with respect to `sigma^2` gives `-n/(2*sigma^2) + (1/(2*sigma^4)) * sum((xi - mu)^2)`; setting that to zero and substituting `mu-hat` gives `sigma^2-hat = (1/n) * sum((xi - xbar)^2)`. The divisor is `n` because that is what the score equation returns; the familiar `n-1` divisor answers a different question about the estimator's sampling behaviour, not the maximisation.

go deeper

for a junior

Recall the two answers: the fitted mean is the sample average and the fitted variance is the average squared distance from it. Being able to compute both on a four-number sample is enough at this level.

for a middle

Be ready to derive both from the Normal log-likelihood at the board, including why the variance cancels out of the mean equation and why substituting the fitted mean is the correct next step.

for a senior

Show that you notice the degenerate cases in practice: a fitted variance collapsing toward zero on near-constant or single-observation groups, and what that says about the data rather than about the algebra.

for a principal

Be able to argue which divisor convention a codebase or reporting standard should use and why the choice is a criterion decision, so that fitted scale parameters stay comparable across teams and pipelines.

## The log-likelihood The Normal density with mean `mu` and variance `sigma^2` is `f(x) = (1 / sqrt(2*pi*sigma^2)) * exp(-(x - mu)^2 / (2*sigma^2))` With `n` independent observations the likelihood is the product of these, and taking logs turns it into `l(mu, sigma^2) = -(n/2)*log(2*pi) - (n/2)*log(sigma^2) - (1/(2*sigma^2)) * sum((xi - mu)^2)` Two unknowns means two score equations, solved one after the other. ## Solving for the mean Only the last term contains `mu`: `d l / d mu = (1/sigma^2) * sum(xi - mu)` Set it to zero. The factor `1/sigma^2` is strictly positive, so it cancels, leaving `sum(xi - mu) = 0`, hence `sum(xi) = n*mu`, hence `mu-hat = xbar` Two things are worth saying out loud. First, the variance disappeared from this equation, so the maximum likelihood mean does not depend on knowing or estimating `sigma^2` -- the two problems partly decouple. Second, minimising `sum((xi - mu)^2)` over `mu` is the same algebra, which is the first hint that squared-error fitting and Normal likelihoods are the same object viewed from different sides. ## Solving for the variance Treat `v = sigma^2` as the parameter (differentiating in `v` rather than `sigma` keeps the algebra clean): `d l / d v = -n/(2*v) + (1/(2*v^2)) * sum((xi - mu)^2)` Setting to zero and multiplying through by `2*v^2` gives `-n*v + sum((xi - mu)^2) = 0`, so `v-hat = (1/n) * sum((xi - mu)^2)` Substituting the already-solved `mu-hat = xbar`: `sigma^2-hat = (1/n) * sum((xi - xbar)^2)` The estimate is the mean squared deviation from the sample mean. ## Why the divisor is n here Candidates who have memorised "the sample variance divides by n-1" often reflexively write `n-1` and are then unable to say where it came from. The honest answer is that the maximisation does not produce it. Setting the derivative of the Normal log-likelihood to zero returns `n`, full stop -- there is no step in the calculus where an `n-1` can appear. The `n-1` divisor is the answer to a different question, about how the estimator behaves across repeated samples, and it is decided by a different criterion than maximising the likelihood. If you are asked "which is right?", the correct framing is that they optimise different objectives, so both are right about their own objective. On this leaf the point to nail is simply: the likelihood's answer is `n`. ## A worked check Take the four values 2, 4, 4, 6. Then `xbar = 4`, and the squared deviations are 4, 0, 0, 4, summing to 8. The maximum likelihood variance is `8/4 = 2`, and hence the maximum likelihood standard deviation is `sqrt(2) ~ 1.41`. Notice that the second step read the standard deviation straight off the variance estimate rather than re-running the maximisation; that shortcut is the invariance property of maximum likelihood. ## Confirming it is a maximum As a function of `mu` with `v` fixed, the log-likelihood is a downward parabola, so `xbar` is clearly its maximum. As a function of `v` with `mu` fixed at `xbar`, the log-likelihood behaves like `-(n/2)*log(v) - S/(2*v)` where `S = sum((xi - xbar)^2)` -- it tends to minus infinity as `v` approaches 0 and as `v` grows large, so its single stationary point at `v = S/n` is the maximum. ## Degenerate cases If every observation is identical then `S = 0` and the maximiser sits at `sigma^2-hat = 0`, a boundary point where the Normal density is not defined; the likelihood is unbounded there. With `n = 1` the same thing happens, since a single point has zero deviation from its own mean. These are the standard reminders that a variance estimate needs more than one distinct observation to mean anything, and that maximum likelihood can run off the edge of the parameter space when the model is asked for more than the data contain. ## What interviewers listen for A correct pair of estimates, the observation that `sigma^2` cancels out of the mean equation, and a calm, accurate explanation of the `n` divisor rather than a panicked switch to `n-1`. Getting that last part right signals that you derived the result rather than recalled it.

  • Why does the mean estimate not depend on sigma-squared?
    Because the derivative with respect to `mu` is `(1/sigma^2) * sum(xi - mu)`, and `1/sigma^2` is a strictly positive factor that cancels when you set the expression to zero. Whatever the variance is, the equation reduces to `sum(xi - mu) = 0`, so `mu-hat = xbar`. The two score equations partly decouple, which is why the mean can be solved first.
  • Compute the maximum likelihood variance for the sample 2, 4, 4, 6.
    The sample mean is 4. The squared deviations are 4, 0, 0 and 4, summing to 8. Dividing by `n = 4` gives `sigma^2-hat = 2`, so the fitted standard deviation is about 1.41. The divisor is the sample size because that is what setting the score equation to zero returns.
  • What happens to the maximum likelihood fit when all observations are identical?
    The sum of squared deviations is zero, so the variance estimate is driven to zero, which is the boundary of the parameter space where the Normal density is undefined and the likelihood is unbounded. The same degeneracy appears with a single observation. It signals that the data cannot support a scale parameter, not that the algebra failed.

saying these in an interview costs you the question

  • Writes n-1 as the divisor the maximisation produced
  • Claims the mean estimate requires knowing sigma-squared first
  • Forgets to substitute the fitted mean into the variance equation
  • Differentiates with respect to sigma and drops a factor
  • Cannot state the estimates for a four-number sample

context