skip to content

When would you refuse to trust the asymptotic standard errors and tests from a likelihood fit?

level: principalimportance: nice to knowfreq 28%

answer

  1. it is a limit result, not an identity
  2. count information, not rows
  3. estimates pinned at an edge
  4. flat directions invert badly
  5. curvature equals score variance only if the model is right

basics

~20 s

Distrust them when the sample is small relative to the number of parameters, when an estimate sits on a boundary of its parameter space, when the model is barely identified, or when the model is misspecified. Each breaks a precondition behind the asymptotics.

solid answer

~50 s

The asymptotic apparatus — Normal sampling distributions, inverse-information standard errors, chi-square reference distributions for nested comparisons — rests on preconditions, and I check them before quoting numbers. First, sample size relative to parameter count: the results are limits, and a fit with few effective observations per parameter can be badly calibrated. Second, boundaries: an estimate at the edge of its space invalidates the two-sided quadratic expansion, so symmetric intervals and the usual chi-square reference are wrong there. Third, identifiability: a near-singular information matrix means the data cannot separate some parameter combination, and its inverse produces huge, unstable errors. Fourth, misspecification: when the model is wrong, the variance of the score and the expected curvature no longer coincide, so nominal errors understate uncertainty and a sandwich estimator is the honest reply. When one of these bites, I move to likelihood-based intervals, resampling calibration, or a reparameterisation rather than reporting the nominal numbers with a footnote.

go deeper

for a junior

Recall that these standard errors and tests are large-sample approximations, so they can be unreliable when data are scarce. Knowing they are not exact is the expected takeaway here.

for a middle

Be ready to name at least two concrete preconditions and what breaks when each fails, for example a boundary estimate invalidating a symmetric interval or a flat direction inflating errors.

for a senior

Show the diagnostic loop: inspect the information matrix conditioning, check for boundary estimates, compare likelihood-ratio against Wald conclusions, and switch to resampled or likelihood-based intervals when they diverge.

for a principal

Own the standard the organisation reports to. Decide when nominal errors are acceptable, mandate robust or resampled variances under dependence, and make sure the failure checks happen before results reach a decision rather than after.

## Why this is a judgment question Every likelihood-based standard error, Wald test and nested-model comparison is the payoff of one approximation: near its maximum, the log-likelihood looks like a quadratic bowl, and the estimate is approximately Normal with variance given by the inverse information. This is a **limit** statement. The senior skill is not reciting it but knowing the situations where the approximation is thin enough that reporting the numbers would mislead a decision. ## Failure mode 1: too little data per parameter Asymptotic results describe behaviour as the sample grows with the model held fixed. A rich model fitted to a modest dataset is exactly the opposite regime. Symptoms: intervals of the form estimate plus or minus two standard errors under-cover; test statistics are too liberal; the fit moves substantially when a few observations are removed. What matters is not raw row count but *effective* information — in a rare-event setting, thousands of rows carrying a handful of events give the information of a very small sample, and the parameter governing those events is the one whose errors are least trustworthy. **Response**: prefer intervals obtained by inverting the likelihood-ratio statistic, which follow the actual shape of the log-likelihood, or calibrate by resampling. Simplify the model until the information supports it. ## Failure mode 2: boundary estimates The quadratic expansion needs the estimate to sit in the **interior** of the parameter space so the log-likelihood can curve away in both directions. When a variance component is estimated at zero, a probability at one, or a mixing weight at an endpoint, the surface is one-sided. Consequences: a symmetric interval can extend into impossible values; the observed information may be singular or ill-defined; and the nested-model comparison no longer has its usual chi-square reference — for a single non-negative variance component tested at zero the correct null distribution is a 50:50 mixture of a point mass at zero and a chi-square with 1 degree of freedom, so using the plain chi-square is conservative rather than merely approximate. **Response**: reparameterise so the boundary is pushed to infinity where that is meaningful, use the correct boundary reference distribution, or move to an interval that respects the parameter's range. ## Failure mode 3: weak identifiability If some combination of parameters barely affects the fit, the log-likelihood has a nearly flat ridge, the information matrix has an eigenvalue near zero, and its inverse explodes. You see enormous standard errors, correlations near plus or minus one between estimates, results that swing with the optimiser's starting point, and a Hessian that fails to be negative definite at the reported solution. The numbers are not merely imprecise — they are numerically unreliable, and small changes in the data or the algorithm change them a lot. **Response**: diagnose rather than report. Drop or combine the offending parameters, add a penalty or prior that regularises the flat direction, or design data collection that varies along it. Never present the raw inverse of a near-singular matrix as an uncertainty statement. ## Failure mode 4: misspecification The identity between the variance of the score and the negative expected curvature — the information equality — holds only when the model is right. Under misspecification the estimate still converges to a well-defined quantity (the parameter minimising divergence from the truth), but its variance is no longer the inverse information. The consequence is systematic: with dependence such as clustered or serially correlated observations, nominal errors are typically **too small**, so tests reject too often. **Response**: a sandwich variance estimator that combines the score's empirical variance with the curvature, or a resampling scheme that respects the dependence structure — resampling clusters, not rows. ## Failure mode 5: statistics that are not invariant Even with plenty of data, a Wald-style test built on the estimate and its standard error depends on how the parameter is written; testing a coefficient and testing its exponential are not the same test. In binary-outcome models the Wald statistic can even *shrink* as the estimate moves further from the null value, an anomaly known as the Hauck-Donner effect, so a very strong effect can produce an unimpressively small statistic. The likelihood-ratio version is invariant to reparameterisation and does not suffer this. **Response**: prefer likelihood-ratio-based tests and intervals when the two disagree, and treat a disagreement between them as a signal, not a rounding difference. ## Turning this into practice As the person responsible for what the organisation believes, the useful move is a standing checklist attached to any likelihood-based result: how much effective information per parameter, is any estimate at a boundary, what is the condition of the information matrix, what dependence structure does the data have, and do the likelihood-ratio and Wald conclusions agree. Where the checklist trips, the report should carry the interval that survives scrutiny rather than the default one plus a caveat nobody reads. The deeper point to convey in an interview is that asymptotic likelihood theory is a superb default precisely because its failure modes are enumerable — and someone has to own the enumeration.

  • What specifically goes wrong when a variance parameter is estimated at zero?
    The estimate sits on the boundary of its space, so the log-likelihood cannot curve away in both directions and the quadratic approximation fails on one side. Symmetric intervals may cover negative variances, the information can be singular, and the nested-model comparison loses its plain chi-square reference — for one such component the correct null is a 50:50 mixture of a point mass at zero and chi-square with 1 degree of freedom.
  • How does clustered or correlated data affect nominal likelihood standard errors?
    A likelihood written as if observations were independent overstates the effective sample size, so the reported errors are typically too small and tests reject too often. The fix is a variance estimator that combines the empirical variability of the score with the model's curvature, or resampling at the cluster level so the dependence is preserved in each resample.
  • How would you decide between reporting Wald and likelihood-ratio intervals as a team default?
    Make likelihood-ratio the default where fitting twice is affordable: it is parameterisation-invariant, follows the actual shape of the log-likelihood, and behaves better in small samples and near boundaries. Reserve Wald intervals for large, well-identified fits where refitting is expensive, and require that any disagreement between the two be investigated rather than averaged over.

saying these in an interview costs you the question

  • Treats asymptotic errors as valid at any sample size
  • Reports symmetric intervals for a parameter pinned at its boundary
  • Inverts a near-singular information matrix and quotes the result
  • Assumes nominal errors are correct under clustered data
  • Judges sample adequacy by row count rather than effective information
  • Treats Wald and likelihood-ratio disagreement as a rounding difference

context