What is Fisher information, and how does it relate to the score function of a log-likelihood?
answer
- start from the log-likelihood's derivatives
- the first derivative is a random variable
- its mean is zero at the truth
- variance of that, or negative expected curvature
- sharper peak means more of it
basics
~20 sThe score is the derivative of the log-likelihood with respect to the parameter; Fisher information is the variance of the score at the true parameter, equivalently minus the expected second derivative. It measures how sharply data pin down the parameter.
solid answer
~40 sFor a parameter `theta`, the score is `U(theta) = d/dtheta log L(theta)`. Under the usual regularity conditions the score has mean zero at the true parameter, `E[U(theta)] = 0`, and Fisher information is its variance: `I(theta) = Var(U(theta)) = E[U(theta)^2]`. The same quantity equals the expected curvature of the log-likelihood, `I(theta) = -E[d^2/dtheta^2 log L(theta)]`, which is the form people actually compute. Intuitively, information is how sharply peaked the log-likelihood is: a steeply curved peak means small changes in `theta` change the fit a lot, so the data are informative. Information adds over independent observations, so `n` i.i.d. draws carry `n * I_1(theta)`. For one Bernoulli observation `I_1(p) = 1 / (p(1-p))`, which is largest near `p = 0` or `p = 1` and smallest at `p = 0.5`.
go deeper
Recall the vocabulary: score is the first derivative of the log-likelihood, information is built from the second. Being able to say that more curvature means more precision is already a good answer at this stage.
Be ready to write both definitions on a whiteboard and show they match, then compute the information for a single Bernoulli observation and note that information adds across independent draws.
Expect to connect information to what you report in practice: where standard errors come from, why nearly flat directions produce huge correlated errors, and what a near-singular information matrix says about identifiability.
Own the misspecification angle: the equality of the two definitions is a modelling assumption, and when it fails you owe the organisation robust variance estimates rather than nominal ones.
## Setting up Suppose data `x_1, ..., x_n` are modelled by a density or mass function `f(x; theta)` indexed by an unknown parameter `theta`. The likelihood `L(theta)` is that model evaluated at the observed data and read as a function of `theta`, and the log-likelihood is `l(theta) = log L(theta)`. Everything in likelihood asymptotics is built from the first two derivatives of `l`. ## The score function The **score** is the first derivative of the log-likelihood: `U(theta) = d/dtheta l(theta)` Setting `U(theta) = 0` is the estimating equation whose solution is the maximum-likelihood estimate, but the score is interesting in its own right as a random quantity: it is a function of the data, so it has a distribution. Under **regularity conditions** — chiefly that the support of `f` does not depend on `theta`, and that you may swap the order of differentiation and integration — the score has mean zero when evaluated at the true parameter: `E[U(theta)] = 0` The one-line derivation: the density integrates to 1 for every `theta`, so differentiating `integral f(x; theta) dx = 1` gives `integral d/dtheta f(x; theta) dx = 0`. Writing `d/dtheta f = f * d/dtheta log f` turns the left side into `E[U(theta)]`, hence zero. ## Fisher information: two equivalent definitions **Definition 1 — variance of the score.** Because the score has mean zero, its variance is just its second moment: `I(theta) = Var(U(theta)) = E[U(theta)^2]` **Definition 2 — expected negative curvature.** Differentiating the mean-zero identity a second time gives the *information equality*: `I(theta) = -E[d^2/dtheta^2 l(theta)]` The second form is usually the easier one to compute, and it is the one that gives information its interpretation. `d^2 l / dtheta^2` is the curvature of the log-likelihood. A sharply peaked log-likelihood (large negative curvature) means that moving `theta` even slightly away from its best value makes the data much less probable — the data strongly discriminate between nearby parameter values. A flat log-likelihood means many parameter values explain the data about equally well, and information is small. Note that the two definitions coincide **only** under the regularity conditions above; the information equality is exactly what fails for misspecified models, which is why robust (sandwich) standard errors combine both quantities instead of using either alone. ## Additivity over observations For independent observations the log-likelihood is a sum, so the score is a sum and — by independence — the information adds: `I_n(theta) = n * I_1(theta)` where `I_1` is the information in a single observation. This additivity is why sample size shows up as a factor of `n` in every downstream result, and ultimately why precision improves like `1/sqrt(n)`. ## Worked example: Bernoulli For one observation `x` in {0, 1} with success probability `p`: - log-likelihood: `l(p) = x log p + (1-x) log(1-p)` - score: `U(p) = x/p - (1-x)/(1-p)` - second derivative: `-x/p^2 - (1-x)/(1-p)^2` - negated expectation, using `E[x] = p`: `p/p^2 + (1-p)/(1-p)^2 = 1/p + 1/(1-p) = 1 / (p(1-p))` So `I_1(p) = 1 / (p(1-p))`. Sanity checks: the score has mean zero, since `E[U(p)] = p/p - (1-p)/(1-p) = 0`. And the information is minimised at `p = 0.5` and blows up as `p` approaches 0 or 1 — a coin that almost always lands heads is highly informative per flip about *how* extreme `p` is, whereas a fair coin's outcomes are maximally ambiguous. With `n` independent flips, total information is `n / (p(1-p))`, and its inverse `p(1-p)/n` is the asymptotic variance attached to the estimate — the precision statement falls straight out of `1 / sqrt(n * I_1)`. ## Multiple parameters When `theta` is a vector, the score is a gradient vector and the information is a matrix: `I(theta) = E[U U^T] = -E[H]`, where `H` is the Hessian of the log-likelihood. Off-diagonal entries record how much two parameters trade off against each other. Large off-diagonal information means the parameters are hard to separate — the log-likelihood ridge runs diagonally, and each parameter individually is poorly determined even when their combination is well determined. ## Why interviewers ask Score and information are the two objects everything else in likelihood theory is written in: the precision bound on unbiased estimators, the asymptotic distribution of the maximum-likelihood estimate, the standard errors reported next to fitted coefficients, and the three classical test statistics. A candidate who can define the score, state the mean-zero property, give both forms of information, and explain the curvature intuition has the vocabulary for all of it.
- Why does the score have expectation zero at the true parameter?Because the density integrates to one for every parameter value. Differentiating that identity under the integral sign gives `integral d/dtheta f(x; theta) dx = 0`, and rewriting `d/dtheta f` as `f * d/dtheta log f` turns the integral into `E[U(theta)]`. So the mean-zero property is a consequence of the regularity conditions that allow that interchange, not a separate assumption.
- When do the two definitions of Fisher information disagree?They agree only when the model is correctly specified and the usual regularity conditions hold. Under misspecification the variance of the score and the negative expected Hessian differ, and their mismatch is exactly what sandwich (robust) variance estimators correct for by combining both. Support that depends on the parameter, such as a uniform on (0, theta), breaks the derivation entirely.
- What does a near-zero eigenvalue of the information matrix tell you?That some linear combination of parameters is nearly unidentified: the log-likelihood is almost flat along that direction, so the data barely distinguish those parameter values. Practically it shows up as huge, strongly correlated standard errors and an unstable fit. The fix is reparameterising, adding constraints or a penalty, or collecting data that varies along the flat direction.
Think of the log-likelihood as a hill whose summit is the best-fitting parameter. The score is the slope you feel underfoot; information is how sharply the summit is pinched. A needle-sharp peak locates the top precisely, a broad plateau leaves you unsure where you are.
saying these in an interview costs you the question
- Calls the score the log-likelihood itself rather than its derivative
- Says Fisher information is the variance of the estimator
- Claims the score has mean zero at every parameter value, not just the truth
- Uses plus the expected second derivative, dropping the minus sign
- Thinks information is a property of the data alone, not of the model and parameter