A row with huge Cook's distance dominates your regression - how do you decide whether to drop it?
answer
- diagnostic locates, it never convicts
- check the record before touching the fit
- typo and genuine extreme differ completely
- population scope is a stated rule
- report the fit both ways
basics
~20 sFirst establish what the row actually is. Delete it only if it is a data error or falls outside the population you are modelling. If it is genuine, keep it and report the fit both with and without it.
solid answer
~50 sThe diagnostic tells you a row matters; it never tells you the row is wrong. So I go and look at it. A ten-times data-entry typo on one revenue field is a defect: correct it from source if the true value is recoverable, otherwise drop it and say so. A genuine extreme enterprise customer is a different case entirely — it is real, it is in the population I was asked about, and deleting it would be quietly changing the question to "what does the model say about everyone except the customers who matter most?". There I keep it and report both fits, stating how much the coefficient moves. The third possibility is that the model is at fault: a missing predictor, a wrong functional form, or an unlogged heavy-tailed response can manufacture an influential point that dissolves once the specification is right.
go deeper
Be ready to say you would look at the actual record first, and that a big influence value is a reason to investigate rather than a licence to delete.
Explain the branch points: data error versus genuine extreme versus wrong model specification, and what each one implies for the fit you report.
Demonstrate lived judgment. Talk through verifying the record, refitting with an alternative specification, and delivering a sensitivity analysis instead of a single tidied-up estimate.
Own the policy. Argue for exclusion rules written before the outcome is seen, uniform application across analyses, and a standard that any conclusion resting on a few rows must say so in the deliverable.
## The rule: diagnose before you delete An influence diagnostic is a smoke alarm. It tells you where to look; it does not tell you what is burning. The single worst answer to this question is "drop it and refit", and the second worst is "keep everything, deleting data is cheating". Both skip the only step that matters — finding out what the row actually represents. ## Step 1: identify the record Go back to the source. Is this a real entity? Does the value pass a plausibility check against other fields on the same record? A revenue figure exactly ten times its neighbours, a height recorded in centimetres in a column of metres, a date that puts an account before the company existed — these are **defects**, and the influence diagnostic simply found one for you. Correct the value if the true one is recoverable; if it is not, remove the row and record the removal and the reason. ## Step 2: ask whether it belongs to the target population A row can be perfectly accurate and still not belong. If you were asked to model small-business accounts and one row is a national government contract, the value is real but it is not a draw from the population your conclusion is about. Excluding it is then a **scope decision**, and the honest way to present it is as a stated inclusion criterion applied uniformly — not as "we removed an outlier". The test of good faith is whether the rule you applied would have excluded the row before you saw its effect on the coefficient. ## Step 3: consider that the model is the problem Influence is defined relative to the specification you fitted. A single genuine extreme enterprise customer can look catastrophic in a straight-line fit of revenue on headcount and completely ordinary once revenue is modelled on a log scale, or once the segment indicator that explains it is included, or once a curved relationship replaces the line. Before concluding that the observation is anomalous, refit with the obvious alternative specifications. If the influence disappears, the row was never the anomaly — the functional form was. ## Step 4: if it is real, in scope, and the model is right, keep it and report both ways This is the case people find uncomfortable and it is usually the right answer. A genuine extreme customer is information: it is telling you the relationship is heavy-tailed, or that a single segment drives the outcome. Deleting it produces a tidier fit that answers a question nobody asked. The deliverable is a **sensitivity analysis**: report the coefficient with the row and without it, state the size and direction of the shift, and let the reader see how much of the conclusion rests on one line of data. If the two fits support the same decision, say so and move on. If they do not, that is the headline finding — the analysis is under-determined and no amount of diagnostics will decide it for you. Collecting more observations in that region of the predictor space is the real fix. If you want a middle path, models that down-weight extreme residuals rather than deleting rows exist and are a defensible alternative; the cost is that the estimate now answers a slightly different question, and you must say which one you are reporting. ## What never justifies deletion - The coefficient becomes significant without the row. - The fit's explained variance improves. - The residual plot looks nicer. Deleting rows until diagnostics look clean is a specification search conducted on the response variable. The standard errors from the surviving fit understate the true uncertainty, because they are computed as if the retained rows were the sample you set out to collect. Confidence intervals get too narrow and p-values too small, in a way no reader can detect from the output. ## Pre-commit where you can The strongest version of the answer is procedural: decide the exclusion rules **before** looking at the outcome. Written data-quality rules ("revenue above the plausible ceiling for the plan tier is a suspected typo and goes to review") and written scope rules ("accounts above 500 seats are out of scope for this model") applied uniformly are defensible in a way that per-row judgement after seeing the coefficient never is. When you do have to make a call after the fact, document the row, the reason, and the effect — a reader who disagrees with your call can then reverse it. ## How to say it in an interview Name the three outcomes explicitly — fix, exclude with a stated rule, or keep and report both ways — and say which evidence pushes you to each. That structure, plus the observation that the model may be at fault, is the whole answer.
- What exactly would you put in the write-up if you keep the influential row?The coefficient with and without the row, the size and direction of the shift relative to its standard error, the identity of the row in business terms, and the decision plus its reason. A reader who disagrees with the call should be able to reverse it from what you reported, without rerunning anything.
- How could a missing predictor manufacture an apparently influential point?If a row belongs to a subgroup with a different intercept or slope and no term captures that, the fit must miss it badly, and if the row also sits far out in the predictors it will swing the line. Add the subgroup indicator, or the interaction, and the influence often disappears — the specification was the anomaly, not the observation.
- Why is deleting rows until the diagnostics look clean statistically dishonest?Because the retained sample was chosen using the response variable. The reported standard errors are computed as if that sample had been drawn independently, so intervals come out too narrow and p-values too small. The output gives a reader no way to see that the selection happened, which is what makes it dishonest rather than merely wrong.
saying these in an interview costs you the question
- Deletes the row because the fit improves
- Refuses ever to exclude data on principle
- Never checks the source record
- Ignores that the model may be misspecified
- Excludes rows without documenting the rule