skip to content

Why prefer studentised deleted residuals over standardised residuals when flagging regression outliers?

level: middleimportance: nice to knowfreq 28%

answer

  1. the suspect inflates its own yardstick
  2. leave-one-out error estimate
  3. raw residual variance depends on leverage
  4. gives a genuine t reference distribution
  5. degrees of freedom lose one more

basics

~20 s

A standardised residual divides by an error scale that the suspect observation itself inflates, which can mask it. The studentised deleted version re-estimates that scale with the observation left out, so a genuine outlier cannot hide behind its own effect on the fit.

solid answer

~50 s

The raw residual is a poor yardstick because its variance is not constant: `Var(e_i) = sigma^2 (1 - h_ii)`, so high-leverage rows have systematically small residuals. Dividing by `s * sqrt(1 - h_ii)` fixes the leverage part and gives the standardised (internally studentised) residual `r_i`. The remaining problem is `s`, the error estimate from the full fit, which the suspect row itself inflates — a bad enough point can enlarge the ruler it is being measured with. The deleted, or externally studentised, residual replaces `s` with `s_(i)`, computed from a fit that excludes row `i`. That version has a clean reference distribution under normal errors: `t` with `n - p - 1` degrees of freedom, where `p` counts the estimated coefficients. So you can judge "how extreme is this" instead of eyeballing a scale-free number.

go deeper

for a junior

Be ready to say residuals are rescaled before being compared, and that the deleted version leaves the suspect observation out when estimating the error scale.

for a middle

Explain the mechanics: dividing by s times sqrt(1 - h_ii), why the full-fit s is circular, and that the deleted version follows a t distribution with n minus p minus 1 degrees of freedom.

for a senior

Show judgment about use. Interviewers want to hear that the two rank rows identically, that the gain is calibration, and that scanning every row needs a multiplicity adjustment before you act on a flag.

for a principal

Own what gets standardised across the team: which residual definition appears in reports, what cutoff triggers investigation, and why a calibrated statistic beats an eyeballed threshold when analysts must defend exclusions.

## Three residuals, one row For observation `i` in an OLS fit with `n` rows and `p` estimated coefficients, there are three quantities people loosely call "the residual". **Raw residual.** `e_i = y_i - yhat_i`. Simple, in the units of the response, and treacherous as a comparison device, because its variance is not the same for every row: ``` Var(e_i) = sigma^2 (1 - h_ii) ``` A row with high leverage `h_ii` drags the fitted surface toward itself, so its raw residual is squeezed toward zero. Ranking rows by raw residual therefore systematically under-ranks exactly the rows most able to distort the fit. **Standardised (internally studentised) residual.** ``` r_i = e_i / (s * sqrt(1 - h_ii)) ``` where `s` is the residual standard error from the full fit. The `sqrt(1 - h_ii)` factor corrects the leverage-induced shrinkage, which is genuine progress: now all the `r_i` have roughly comparable spread. People commonly eyeball `|r_i| > 2` or `> 3` as worth a look. **Studentised deleted (externally studentised) residual.** ``` t_i = e_i / (s_(i) * sqrt(1 - h_ii)) ``` where `s_(i)` is the residual standard error from a model estimated *without* observation `i`. Everything is the same except the ruler. ## Why the ruler matters The defect in `r_i` is circular reasoning. `s` is computed from all the residuals including `e_i`. A badly discrepant observation therefore inflates `s`, and inflating `s` shrinks `r_i` — the point pushes down its own measured extremeness. With a single bad row in a small sample this can be enough to keep it under an informal cutoff. Using `s_(i)`, an estimate the suspect row never touched, removes the circularity: the observation is judged against how well the *rest* of the data are fitted. This is the same leave-one-out logic that underlies influence diagnostics generally. Ask what the model would look like if this row had never existed, then ask how surprising the row is under that model. ## The reference distribution The practical payoff is calibration. Under the standard assumption of independent normal errors, the studentised deleted residual follows a **t distribution with `n - p - 1` degrees of freedom**. The arithmetic is direct: the deleted fit uses `n - 1` observations and `p` parameters, leaving `n - 1 - p` degrees of freedom for its error estimate. The internally studentised `r_i` has no such clean distribution — it is a bounded quantity, not a t variate, which is why the `|r| > 2` habit is a rule of thumb rather than a test. One caution about using that distribution as a test. You are typically not asking about one pre-chosen row; you are scanning all `n` rows and reacting to the largest. That is a multiple-comparison problem, and the usual fix is a Bonferroni adjustment: compare the largest `|t_i|` against the `alpha/(2n)` tail of the t distribution rather than the `alpha/2` tail. Without that correction, a few rows past `|t| = 3` in a dataset of ten thousand is exactly what normal errors would produce anyway. ## They rank the same, so why bother? A fair objection: `t_i` is a monotone increasing function of `r_i`, ``` t_i = r_i * sqrt((n - p - 1) / (n - p - r_i^2)) ``` so the two orderings of the rows are identical. Nothing changes about *which* row is most extreme. What changes is the interpretation of the number. The internally studentised residual is bounded by construction and has no reference distribution, so a value of 3.4 means whatever your habit says it means. The deleted version can be compared to a t distribution and, with the multiplicity adjustment, converted into a defensible statement about how surprising the most extreme row is. When the transformation blows up — `r_i^2` approaching `n - p` — that itself is the signal of one dominating row. ## How it connects to influence Studentised residuals are the discrepancy half of influence, not influence itself. Combine one with the leverage and you get an influence measure: DFFITS, for instance, is `t_i * sqrt(h_ii / (1 - h_ii))`, the deleted residual multiplied by the lever-arm factor. So the correct reading is that a large `|t_i|` says "this response disagrees with what the rest of the data imply", and only pairing it with leverage tells you whether that disagreement moved the fit. ## The short interview version Raw residuals are unfairly small for high-leverage rows; standardised residuals fix that but are measured against a ruler the suspect row helped build; deleted residuals fix the ruler too and come with a `t_{n-p-1}` reference distribution, at the cost of needing the leave-one-out error estimate — which has a closed form, so it costs nothing in practice.

  • Since the deleted residual is a monotone function of the standardised one, what do you actually gain?
    Not a different ranking — the ordering of the rows is identical. What you gain is calibration: the deleted version follows a t distribution with `n - p - 1` degrees of freedom under normal errors, so its value is interpretable as a tail probability, whereas the internally studentised residual is a bounded quantity with no clean reference distribution.
  • How would you avoid over-flagging when scanning every row's studentised residual?
    Treat it as a multiplicity problem. If you react to the largest of `n` values, compare it against the `alpha/(2n)` tail of the t distribution rather than the `alpha/2` tail. On large data a handful of values past three is exactly what well-behaved normal errors produce, so an uncorrected scan generates constant false alarms.
  • Does a large studentised residual by itself mean the row changed the fit?
    No. It says the response disagrees with what the rest of the data imply, which is discrepancy, not influence. A discrepant row near the centre of the predictor range has a short lever arm and moves the slope very little. You need to pair the residual with the leverage before claiming the fit was affected.

Judging how unusual a person's height is with a tape measure that the person themselves stretched. Leave them out while calibrating the tape, and the measurement finally means something.

saying these in an interview costs you the question

  • Thinks the two residual types rank rows differently
  • Uses the full-fit error estimate to judge the suspect row
  • Quotes n minus one degrees of freedom
  • Treats a large studentised residual as proof of influence
  • Scans thousands of rows with no multiplicity adjustment

context