Why prefer studentised deleted residuals over standardised residuals when flagging regression outliers?
answer
- the suspect inflates its own yardstick
- leave-one-out error estimate
- raw residual variance depends on leverage
- gives a genuine t reference distribution
- degrees of freedom lose one more
basics
~20 sA standardised residual divides by an error scale that the suspect observation itself inflates, which can mask it. The studentised deleted version re-estimates that scale with the observation left out, so a genuine outlier cannot hide behind its own effect on the fit.
solid answer
~50 sThe raw residual is a poor yardstick because its variance is not constant: `Var(e_i) = sigma^2 (1 - h_ii)`, so high-leverage rows have systematically small residuals. Dividing by `s * sqrt(1 - h_ii)` fixes the leverage part and gives the standardised (internally studentised) residual `r_i`. The remaining problem is `s`, the error estimate from the full fit, which the suspect row itself inflates — a bad enough point can enlarge the ruler it is being measured with. The deleted, or externally studentised, residual replaces `s` with `s_(i)`, computed from a fit that excludes row `i`. That version has a clean reference distribution under normal errors: `t` with `n - p - 1` degrees of freedom, where `p` counts the estimated coefficients. So you can judge "how extreme is this" instead of eyeballing a scale-free number.
go deeper
Be ready to say residuals are rescaled before being compared, and that the deleted version leaves the suspect observation out when estimating the error scale.
Explain the mechanics: dividing by s times sqrt(1 - h_ii), why the full-fit s is circular, and that the deleted version follows a t distribution with n minus p minus 1 degrees of freedom.
Show judgment about use. Interviewers want to hear that the two rank rows identically, that the gain is calibration, and that scanning every row needs a multiplicity adjustment before you act on a flag.
Own what gets standardised across the team: which residual definition appears in reports, what cutoff triggers investigation, and why a calibrated statistic beats an eyeballed threshold when analysts must defend exclusions.
## Three residuals, one row For observation `i` in an OLS fit with `n` rows and `p` estimated coefficients, there are three quantities people loosely call "the residual". **Raw residual.** `e_i = y_i - yhat_i`. Simple, in the units of the response, and treacherous as a comparison device, because its variance is not the same for every row: ``` Var(e_i) = sigma^2 (1 - h_ii) ``` A row with high leverage `h_ii` drags the fitted surface toward itself, so its raw residual is squeezed toward zero. Ranking rows by raw residual therefore systematically under-ranks exactly the rows most able to distort the fit. **Standardised (internally studentised) residual.** ``` r_i = e_i / (s * sqrt(1 - h_ii)) ``` where `s` is the residual standard error from the full fit. The `sqrt(1 - h_ii)` factor corrects the leverage-induced shrinkage, which is genuine progress: now all the `r_i` have roughly comparable spread. People commonly eyeball `|r_i| > 2` or `> 3` as worth a look. **Studentised deleted (externally studentised) residual.** ``` t_i = e_i / (s_(i) * sqrt(1 - h_ii)) ``` where `s_(i)` is the residual standard error from a model estimated *without* observation `i`. Everything is the same except the ruler. ## Why the ruler matters The defect in `r_i` is circular reasoning. `s` is computed from all the residuals including `e_i`. A badly discrepant observation therefore inflates `s`, and inflating `s` shrinks `r_i` — the point pushes down its own measured extremeness. With a single bad row in a small sample this can be enough to keep it under an informal cutoff. Using `s_(i)`, an estimate the suspect row never touched, removes the circularity: the observation is judged against how well the *rest* of the data are fitted. This is the same leave-one-out logic that underlies influence diagnostics generally. Ask what the model would look like if this row had never existed, then ask how surprising the row is under that model. ## The reference distribution The practical payoff is calibration. Under the standard assumption of independent normal errors, the studentised deleted residual follows a **t distribution with `n - p - 1` degrees of freedom**. The arithmetic is direct: the deleted fit uses `n - 1` observations and `p` parameters, leaving `n - 1 - p` degrees of freedom for its error estimate. The internally studentised `r_i` has no such clean distribution — it is a bounded quantity, not a t variate, which is why the `|r| > 2` habit is a rule of thumb rather than a test. One caution about using that distribution as a test. You are typically not asking about one pre-chosen row; you are scanning all `n` rows and reacting to the largest. That is a multiple-comparison problem, and the usual fix is a Bonferroni adjustment: compare the largest `|t_i|` against the `alpha/(2n)` tail of the t distribution rather than the `alpha/2` tail. Without that correction, a few rows past `|t| = 3` in a dataset of ten thousand is exactly what normal errors would produce anyway. ## They rank the same, so why bother? A fair objection: `t_i` is a monotone increasing function of `r_i`, ``` t_i = r_i * sqrt((n - p - 1) / (n - p - r_i^2)) ``` so the two orderings of the rows are identical. Nothing changes about *which* row is most extreme. What changes is the interpretation of the number. The internally studentised residual is bounded by construction and has no reference distribution, so a value of 3.4 means whatever your habit says it means. The deleted version can be compared to a t distribution and, with the multiplicity adjustment, converted into a defensible statement about how surprising the most extreme row is. When the transformation blows up — `r_i^2` approaching `n - p` — that itself is the signal of one dominating row. ## How it connects to influence Studentised residuals are the discrepancy half of influence, not influence itself. Combine one with the leverage and you get an influence measure: DFFITS, for instance, is `t_i * sqrt(h_ii / (1 - h_ii))`, the deleted residual multiplied by the lever-arm factor. So the correct reading is that a large `|t_i|` says "this response disagrees with what the rest of the data imply", and only pairing it with leverage tells you whether that disagreement moved the fit. ## The short interview version Raw residuals are unfairly small for high-leverage rows; standardised residuals fix that but are measured against a ruler the suspect row helped build; deleted residuals fix the ruler too and come with a `t_{n-p-1}` reference distribution, at the cost of needing the leave-one-out error estimate — which has a closed form, so it costs nothing in practice.
- Since the deleted residual is a monotone function of the standardised one, what do you actually gain?Not a different ranking — the ordering of the rows is identical. What you gain is calibration: the deleted version follows a t distribution with `n - p - 1` degrees of freedom under normal errors, so its value is interpretable as a tail probability, whereas the internally studentised residual is a bounded quantity with no clean reference distribution.
- How would you avoid over-flagging when scanning every row's studentised residual?Treat it as a multiplicity problem. If you react to the largest of `n` values, compare it against the `alpha/(2n)` tail of the t distribution rather than the `alpha/2` tail. On large data a handful of values past three is exactly what well-behaved normal errors produce, so an uncorrected scan generates constant false alarms.
- Does a large studentised residual by itself mean the row changed the fit?No. It says the response disagrees with what the rest of the data imply, which is discrepancy, not influence. A discrepant row near the centre of the predictor range has a short lever arm and moves the slope very little. You need to pair the residual with the leverage before claiming the fit was affected.
Judging how unusual a person's height is with a tape measure that the person themselves stretched. Leave them out while calibrating the tape, and the measurement finally means something.
saying these in an interview costs you the question
- Thinks the two residual types rank rows differently
- Uses the full-fit error estimate to judge the suspect row
- Quotes n minus one degrees of freedom
- Treats a large studentised residual as proof of influence
- Scans thousands of rows with no multiplicity adjustment