How do you get a standard error for a maximum-likelihood estimate from the log-likelihood?
answer
- precision lives in the curvature at the peak
- second derivatives of the log-likelihood
- negate, then invert the matrix
- square roots of the diagonal
- observed at the data, expected under the model
basics
~20 sTake the Hessian of the log-likelihood at the optimum, negate it to get the observed information matrix, invert it, and read standard errors as the square roots of its diagonal. Curvature at the peak is what precision means here.
solid answer
~50 sMaximum-likelihood standard errors come from curvature. Compute the matrix of second derivatives of the log-likelihood at the estimate, negate it to get the observed information `J(theta_hat)`, invert, and take square roots of the diagonal: `SE(theta_hat_j) = sqrt([J(theta_hat)^-1]_jj)`. This is a direct consequence of asymptotic normality: `sqrt(n) * (theta_hat - theta)` converges to a Normal with variance `1/I_1(theta)`, so the estimate behaves like a draw from `Normal(theta, 1/I_n(theta))` in large samples. You have two choices of information — **observed**, the negated Hessian at the data you actually saw, and **expected**, its expectation under the model with the unknown parameter plugged in. They agree asymptotically; observed information is usually preferred because it reflects this dataset and needs no extra expectation. Off-diagonal entries of the inverse matter too: they are the covariances between parameter estimates, and you need them for any contrast or derived quantity.
go deeper
Recall that the uncertainty of a fitted parameter comes from how sharply the log-likelihood peaks, and that the standard error is the square root of an inverse-curvature quantity.
Be ready to walk the recipe end to end: Hessian at the optimum, negate, invert, square-root the diagonal, and say why that matches the asymptotic Normal result.
Demonstrate judgment about when these errors mislead: small samples, boundary estimates, near-singular information, and contrasts where the covariance term must be carried through.
Own the reporting standard. Decide when the team quotes Wald-style errors versus likelihood-based or resampled intervals, and make sure boundary and identifiability failures are caught before numbers reach a decision-maker.
## The result standard errors rest on Under regularity conditions, the maximum-likelihood estimate is **asymptotically Normal**: `sqrt(n) * (theta_hat - theta) -> Normal(0, 1 / I_1(theta))` Rewriting for practical use, in large samples `theta_hat ~approx~ Normal(theta, 1 / I_n(theta))`, with `I_n(theta) = n * I_1(theta)` So the asymptotic variance of the estimate is the **inverse of the Fisher information in the sample**, and its square root is the standard error. Everything below is machinery for evaluating that inverse when you do not know `theta`. ## Observed versus expected information Let `l(theta)` be the full-sample log-likelihood and `H(theta)` its Hessian (matrix of second partial derivatives). - **Observed information**: `J(theta_hat) = -H(theta_hat)`. Purely a function of the data at hand; no expectation is taken. - **Expected information**: `I(theta) = -E[H(theta)]`, evaluated at `theta_hat` because the true value is unknown. Both are consistent estimates of the same limiting quantity, so they agree as `n` grows. In finite samples the observed version is generally preferred: it requires no analytic expectation (often unavailable), it is what an optimiser already has at the point of convergence, and it conditions on the dataset you actually observed rather than averaging over datasets you did not. The expected version can be smoother and is occasionally preferred when the observed Hessian is badly behaved at the optimum. ## The recipe 1. Maximise the log-likelihood to get `theta_hat`. 2. Evaluate the Hessian there and negate it: `J = -H(theta_hat)`. 3. Invert: `V = J^-1`. This is the estimated asymptotic covariance matrix. 4. Standard errors are `SE_j = sqrt(V_jj)`. 5. Correlations between estimates come from the off-diagonals: `corr(j, k) = V_jk / sqrt(V_jj * V_kk)`. A scalar example makes the shape clear. With `n` Bernoulli trials, the information in the sample is `n / (p(1-p))`, so its inverse is `p(1-p)/n` and the standard error is `sqrt(p_hat(1-p_hat)/n)` after plugging in the estimate. The general recipe reproduces the familiar formula rather than competing with it. ## Interpreting the geometry A sharply curved log-likelihood peak inverts to a small variance: the data strongly prefer one parameter value. A shallow peak inverts to a large variance. When two parameters trade off — a long diagonal ridge in the likelihood surface — the information matrix is near-singular, its inverse has large entries, and the standard errors come out enormous and strongly correlated. That pattern is a diagnostic, not a nuisance: it says the data cannot separate those parameters, and no amount of computation fixes it. ## Functions of parameters: the delta method You often want a standard error for `g(theta_hat)` rather than `theta_hat` itself — an odds ratio, a predicted rate, a ratio of two coefficients. The delta method propagates the covariance through the gradient: `Var(g(theta_hat)) ~approx~ grad_g(theta_hat)^T * V * grad_g(theta_hat)` This is where the off-diagonals earn their keep: for a contrast such as `theta_1 - theta_2`, the variance is `V_11 + V_22 - 2*V_12`, and ignoring the covariance term can be badly wrong in either direction. ## Where the approximation is thin These are **asymptotic** standard errors, and the label is a warning. - **Small samples.** The Normal approximation to the sampling distribution may be poor, and intervals of the form estimate plus or minus two standard errors can undercover. - **Parameters near a boundary.** If the estimate sits at or near an edge of the parameter space — a variance component at zero, a probability at one — the quadratic approximation to the log-likelihood is wrong on one side, and the symmetric interval may cover impossible values. - **Poorly identified directions.** Near-flat directions make the inverse numerically unstable; a Hessian that is not negative definite at the reported optimum usually means the optimiser did not converge to an interior maximum. - **Skewed likelihoods.** The quadratic (Wald) approximation imposes symmetry the likelihood may not have; intervals built by inverting the likelihood-ratio statistic respect the shape better and can be asymmetric. A senior answer names the recipe, distinguishes observed from expected information, and then says where it stops being trustworthy — that last part is what separates someone who has read the theory from someone who has shipped estimates with error bars attached.
- Why choose observed information over expected information in practice?Observed information needs no analytic expectation, which is often intractable, and it is already available as the Hessian at the optimiser's solution. It also conditions on the dataset you actually have rather than averaging over hypothetical ones, which generally gives better-calibrated finite-sample errors. The two coincide asymptotically, so the choice is about finite-sample behaviour and convenience.
- How do you get a standard error for a nonlinear function of the fitted parameters?Use the delta method: `Var(g(theta_hat))` is approximately `grad_g^T * V * grad_g`, where `V` is the inverse observed information. The gradient is evaluated at the estimate. For a simple contrast like `theta_1 - theta_2` this reduces to `V_11 + V_22 - 2*V_12`, which is why dropping the off-diagonal covariance can badly misstate the uncertainty.
- The reported Hessian at the optimum is not negative definite. What does that tell you?That the reported point is not a well-behaved interior maximum. Common causes are non-convergence, a parameter pinned at a boundary, or a direction in which the likelihood is flat because the model is not identified from this data. Inverting such a matrix yields meaningless or negative variances, so diagnose the fit rather than reporting the numbers.
- When would you prefer a likelihood-ratio interval to an estimate plus or minus two standard errors?When the log-likelihood is visibly asymmetric around the peak, the sample is small, or the parameter is near a boundary. Standard-error intervals impose a symmetric quadratic approximation; inverting the likelihood-ratio statistic traces the actual shape of the log-likelihood, gives asymmetric limits, and respects the parameter's range.
The log-likelihood peak is a bowl. A narrow bowl traps a marble tightly, so you know where the parameter is; a wide shallow bowl lets it wander. Inverting the curvature converts bowl shape into that wandering distance.
saying these in an interview costs you the question
- Inverts the Hessian without negating it, producing negative variances
- Takes reciprocals of the diagonal instead of inverting the matrix
- Ignores off-diagonal covariances when reporting a contrast
- Treats asymptotic standard errors as exact at any sample size
- Reports symmetric intervals for a parameter estimated at a boundary