In a regression fit, what is the difference between an outlier and an influential observation?
answer
- one is about y, one about consequence
- leverage is the bridge between them
- large residual near the centre moves little
- drop the row and refit to test it
- influence roughly multiplies the two
basics
~10 sAn outlier has a large residual: its response sits far from the fitted line. An influential observation is one whose removal visibly changes the fitted coefficients. A point can be either, both, or neither.
solid answer
~40 sAn outlier is unusual in `y` given its `x`: the model predicts one thing and the observation says another, so its residual is large. Influence is about consequence — drop the row, refit, and see whether the coefficients move. The bridge between them is leverage, which measures how unusual the row's predictors are. A large residual at the centre of the predictor range mostly lifts the intercept and inflates the error estimate, while the same discrepancy far out in `x` swings the slope. That is why influence measures multiply the two together: a row generally needs both an odd predictor position and an odd response to dominate a fit. In a wealth-on-education regression, one billionaire's row can carry both — extreme education, extreme wealth — and taking it out visibly rotates the line.
go deeper
Be ready to state the two definitions cleanly: outlier means large residual, influential means the coefficients change when you drop it, and the two are not the same thing.
Explain the mechanism that links them — leverage. Walk the four corners of the leverage-by-residual grid and say which corner actually threatens the slope and why the raw residual can hide it.
Demonstrate that you check influence before reacting. Interviewers want to hear you refit without the row, quantify the coefficient shift, and consider that a missing predictor rather than the observation may be the real problem.
Frame it as a reporting standard: when a conclusion rests on a handful of rows, the team should say so. Argue for sensitivity results being part of the deliverable rather than an ad-hoc reaction to a scary plot.
## Three different questions about one row When someone points at a scatter plot and asks about a suspicious point, they are really asking three separate questions, and interviewers want to hear them separated. 1. **Is it an outlier?** Is the response unusual *given* the predictors — is the residual large? An outlier disagrees with the model. 2. **Does it have leverage?** Are the predictor values unusual, far from the centre of the predictor cloud? Leverage is computed from the predictors alone and says nothing about the response. 3. **Is it influential?** Would the conclusions change if the row were not there? This is the only one of the three that is defined by consequence: refit without the row and compare. Influence is roughly the product of the other two. A row that is unremarkable in `x` and unremarkable in `y` cannot move much. A row that is unusual on one axis alone usually moves little. A row that is unusual on both is the dangerous case. ## The four corners **Low leverage, small residual.** An ordinary observation. It contributes to the fit like everyone else and nothing about it needs a decision. **High leverage, small residual.** A point far out along the predictor axis whose response lands almost exactly on the line the rest of the data imply. Deleting it barely moves the slope — indeed it is the observation that pins the slope down most precisely, because it is the one with the longest lever arm. Removing it typically *widens* the coefficient's confidence interval. This corner is why "high leverage" is not a synonym for "problem". **Low leverage, large residual.** A point near the middle of the predictor range that sits well off the line. It is a genuine outlier: it inflates the residual standard error and nudges the intercept, but the slope is comparatively stable because the point sits near the pivot. It shows up loudly in residual summaries and quietly in influence summaries. **High leverage, large residual.** The row that rotates the line. Imagine regressing wealth on years of education across a national sample and leaving one billionaire in the data. That row may be far out in education *and* orders of magnitude off in wealth. It can single-handedly set the slope, and dropping it moves the estimate enough that the story you tell changes. This is the corner where influence diagnostics earn their keep. ## Why the raw residual misleads you A subtlety that separates a good answer from a great one: the residual of a high-leverage point is *systematically shrunk*, because the fitted surface is pulled toward it. Formally `Var(e_i) = sigma^2 (1 - h_ii)`, so as leverage grows the residual's own spread shrinks toward zero. The most dangerous rows can therefore look innocuous in a raw residual list — they have bent the line enough to hide. That is exactly why residuals are rescaled before being judged, and why influence measures combine residual and leverage rather than looking at either alone. ## Outlier in y, not outlier in the marginal sense Regression outlyingness is *conditional*. A revenue value of ten million may be perfectly ordinary in a fit if the predictors say that account should be huge, and a revenue value of eighty thousand may be a glaring outlier if the predictors say it should be five thousand. Judging the response column on its own, without conditioning on the predictors, answers a different question than regression diagnostics do. ## What to actually do The distinction matters because it changes the response. An outlier with no influence is mostly a note in the write-up and a reason to check whether the error assumptions hold. An influential row is a decision: verify the record, decide whether it belongs to the population you are modelling, and if you keep it, report the fit both with and without it so a reader can see how much the conclusion rests on one line of data. A final warning: influence is defined relative to *this* model. A point that dominates a straight-line fit may be perfectly ordinary once you add the predictor that explains it or fit the relationship on a log scale. Before deciding a row is the problem, check whether the model is.
- Give an example of an influential point that is not an outlier.A row far out in the predictor range whose response is only mildly off the line can still swing the slope, because the long lever arm converts a modest discrepancy into a large rotation. Its residual would never make an outlier list, yet deleting it visibly changes the coefficient — influence is leverage times discrepancy, so a big multiplier compensates for a small one.
- How would you check influence if you did not remember any named diagnostic?Refit with the row removed and compare the coefficients you care about, ideally relative to their standard errors. Every standard influence measure is a scaled shortcut for exactly that leave-one-out comparison, so doing it by hand for a handful of suspicious rows is a legitimate and transparent answer.
- Does a large residual mean the observation is wrong?No. It means the model and the observation disagree, and the model is at least as likely to be at fault. Missing predictors, a wrong functional form, or an unmodelled subgroup all produce large residuals from perfectly correct data. Verify the record, but also ask whether the fit deserves the benefit of the doubt.
saying these in an interview costs you the question
- Treats outlier and influential as synonyms
- Says any point far out in x must be influential
- Judges the response column without conditioning on predictors
- Assumes a large residual means the data are wrong
- Never mentions refitting without the row