How does Cook's distance measure the influence of a single observation on a regression fit?
answer
- how far all predictions move on deletion
- discrepancy multiplied by lever arm
- standardised residual squared times h/(1-h)
- screen at 4/n or at 1
- global fit, not one coefficient
basics
~20 sCook's distance summarises how far all the fitted values move when one observation is dropped and the model is refitted. It combines that row's rescaled residual with its leverage, so a point needs both an odd response and an odd predictor position to score high.
solid answer
~40 sCook's distance for observation `i` is the scaled sum of squared changes in *every* fitted value when row `i` is deleted: `D_i = sum_j (yhat_j - yhat_j(i))^2 / (p * s^2)`, where `p` is the number of estimated coefficients and `s^2` the residual mean square. There is a closed form that avoids the refit: `D_i = (r_i^2 / p) * (h_ii / (1 - h_ii))`, with `r_i` the standardised residual and `h_ii` the leverage. That product is the whole idea — discrepancy times lever arm. Common screening cutoffs are `D_i > 4/n` or `D_i > 1`; both are heuristics, not tests. Cook's distance is deliberately global: it asks whether the *fit as a whole* moved. If you care about one specific coefficient, look at DFBETA for that coefficient instead.
go deeper
Be able to say Cook's distance measures how much the fit changes when one observation is removed, and that a big value means that row is worth looking at rather than deleting.
Explain the closed form as standardised residual squared times h/(1-h), divided by the coefficient count, and be ready to compute the 4/n cutoff for a given sample size.
Show you use it as a triage tool. Interviewers want to hear you follow a flagged list with DFBETA on the coefficient that matters, and that you know single-deletion measures can be masked by pairs of similar points.
Own the reporting standard: which influence diagnostics run by default, what thresholds mean in your data sizes, and how the team communicates that a headline estimate depends on a small number of rows.
## The definition Cook's distance answers a blunt question: if I delete this one row and refit, how much does the model's whole set of predictions move? Write `yhat_j(i)` for the fitted value of observation `j` from a model estimated without observation `i`. Then ``` D_i = sum over j of (yhat_j - yhat_j(i))^2 / (p * s^2) ``` where `p` is the number of estimated coefficients (predictors plus the intercept) and `s^2` is the residual mean square from the full fit. The numerator is a squared distance between two whole vectors of predictions; the denominator turns it into a unitless quantity, so `D_i` is comparable across models and scales. ## The closed form, and why it is the interesting part You never actually have to run `n` refits, because the algebra collapses: ``` D_i = (r_i^2 / p) * (h_ii / (1 - h_ii)) ``` Here `r_i = e_i / (s * sqrt(1 - h_ii))` is the standardised residual and `h_ii` is the leverage. Read the two factors separately: - `r_i^2` is **discrepancy**: how badly the model missed this response. - `h_ii / (1 - h_ii)` is **lever arm**: it is near zero for a central point, equals 1 when `h_ii = 0.5`, and explodes as `h_ii` approaches 1. Influence is the product. This is the formula behind every hand-wavy statement that "a point needs to be unusual in both x and y to dominate a fit". A huge residual in the middle of the predictor cloud is multiplied by a tiny lever arm and scores low; a far-out point sitting on the line has a huge lever arm multiplied by a near-zero residual and also scores low. ## Thresholds Two cutoffs are in common circulation: - `D_i > 4/n` — a sensitive screen that scales with sample size. On 250 rows this flags anything above 0.016, which will typically catch a handful of rows in any real dataset. - `D_i > 1` — a much blunter rule aimed at points that move the fit dramatically. Neither is a significance test, and neither licenses deletion. Their job is to rank rows for inspection. On a large dataset `4/n` will always flag *something*, which is a feature when you are screening and a trap if you treat the flag as a verdict. ## Global influence versus coefficient-specific influence Cook's distance is a single number per row, and it deliberately aggregates over all fitted values. That makes it a good screen and a poor answer to the question people usually care about, which is "did this row change *the coefficient in my headline*?" The coefficient-specific diagnostic is **DFBETA**: for row `i` and coefficient `j`, `DFBETA_ij = b_j - b_j(i)`, the raw change in that one coefficient when the row is deleted. Divided by the standard error of `b_j(i)` it becomes the standardised DFBETAS, screened at roughly `2/sqrt(n)`. The two can disagree in a way that is genuinely useful. Suppose a `4/n` screen flags three rows. You compute DFBETA for the one coefficient the stakeholder actually reads, and find that two of the three barely move it — they were dragging *other* parts of the fit, perhaps a nuisance control — while the third shifts it by most of a standard error. Now you have one row to investigate rather than three, and a defensible statement about which conclusion depends on it. **DFFITS** sits between them: `DFFITS_i = (yhat_i - yhat_i(i)) / (s_(i) * sqrt(h_ii))`, the change in row `i`'s own fitted value, scaled by the leave-one-out error estimate. A common cutoff is `2 * sqrt(p/n)`. It is closely related to Cook's distance by construction, so in practice the two rank rows almost identically; carrying both into an interview is fine, but be able to say why you would reach for DFBETA when the question is about one specific coefficient. ## Limits worth naming All of these are **single-deletion** diagnostics. Two nearly identical extreme rows can mask each other: delete either alone and the other holds the line in place, so both score low despite jointly dominating the fit. Detecting that needs multiple-deletion or iterative approaches, and the practical version is to plot the data rather than trust the ranked list. Also remember that `D_i` is defined relative to the model you fitted. Change the functional form, add the omitted predictor, or move to a log scale, and the influence ranking can change completely. A large Cook's distance is a question, not an answer.
- How does DFBETA differ from Cook's distance, and when would you prefer it?Cook's distance aggregates the shift across all fitted values into one global number, while DFBETA reports the change in a *specific* coefficient when the row is deleted. Prefer DFBETA when the conclusion rests on one coefficient: a row can inflate Cook's distance by disturbing nuisance controls while leaving the headline estimate essentially untouched.
- Why does the 4/n threshold flag more rows as the dataset grows?Because the cutoff shrinks with sample size while the distribution of Cook's distances does not vanish. On large data the rule is a ranking device that will always highlight a tail of rows. Treat it as "look at these first", and lean on the absolute size of the coefficient shift to decide whether anything matters.
- Can two influential observations both show small Cook's distances?Yes — this is masking. Cook's distance is a leave-one-out measure, so two nearly identical extreme rows each look harmless because deleting one leaves the other propping the fit up. Joint influence needs multiple-deletion diagnostics or, more practically, actually looking at the plot before trusting a ranked list.
saying these in an interview costs you the question
- Says Cook's distance uses only the residual
- Treats 4/n as a significance test
- Cannot say what it measures the change in
- Claims a high Cook's distance proves bad data
- Uses it to judge one specific coefficient