skip to content

How can a single extreme point push Pearson's r from near zero to 0.8?

level: seniorimportance: should knowfreq 54%

answer

  1. the statistic is a sum over rows
  2. each term multiplies two deviations
  3. far from both means, both deviations large
  4. one term outweighs all the others
  5. drop each row in turn and recompute

basics

~20 s

Pearson's r sums cross-products of deviations, so one point far from both means can contribute a term that outweighs all the others and dominates the statistic. Plot the scatter and recompute r with each row removed to catch it.

solid answer

~50 s

r sums the products `(xi - xbar) * (yi - ybar)` and divides by the two standard deviations. A point far from both means contributes a term proportional to the product of two large deviations, so with 30 unstructured points and one point far out along a line, that single term can outweigh the other 30 combined and drag r from near 0 to 0.8 or beyond. The lever works in reverse too: one point off the axis of an otherwise tight relationship can collapse a high r toward 0. Diagnosis is mechanical. Look at the scatterplot before quoting r, then recompute r with each row dropped in turn and record the largest swing. If one row moves r materially, report r with and without it and investigate whether that row is a data error or a genuine rare case.

go deeper

for a junior

Know that Pearson's r is fragile to extreme values and that a scatterplot should be looked at before any correlation is quoted. Recall that one point far from the rest can raise or lower r substantially.

for a middle

Explain the mechanism: r sums products of two deviations, so a point far from both means contributes a term that can outweigh every other. Describe leave-one-out recomputation as the concrete check.

for a senior

Demonstrate the full loop -- detect, investigate the row, decide, and disclose. Interviewers want to hear that you report r with and without the point rather than quietly dropping it, and that you distinguish a data error from a genuine rare case.

for a principal

Own the standard: which association numbers ship with a sensitivity check, what happens when a headline metric turns out to rest on one row, and how the team avoids a culture of quietly filtering data until the correlation looks right.

## Why one point has so much leverage Pearson's r for n paired observations is `r = sum_i (xi - xbar) * (yi - ybar) / ((n - 1) * sx * sy)` The numerator is a sum of **products of deviations**. Two things make that sum easy for one row to hijack: 1. **The contribution is a product, not a distance.** A point ten deviations from the mean in x and ten in y contributes roughly a hundred times the term of a typical point sitting one deviation out in each direction. 2. **The sum is unweighted and unbounded.** Nothing caps how much one term can add. Push a single point far enough along any straight line and r converges toward plus or minus 1, no matter what the other n - 1 points look like. With 30 points forming a formless blob (r near 0) and one added point far away along a line through the blob, that point's cross-product can exceed the total from the other 30, and the reported r lands at 0.8 or higher. The same mechanism runs backwards: take 30 points lying tightly along a line, add one point far away perpendicular to it, and r can fall toward 0. Identical data, one row of difference, opposite headline. ## Two different kinds of extreme point It helps to separate them: - A point extreme in **both** variables and consistent with a straight line through the data inflates r. This is the classic high-leverage point. - A point extreme in **one** variable but ordinary in the other, or extreme in a direction the rest of the data does not follow, deflates r. The distinction matters when you go looking: scanning for outliers one column at a time can miss a pair that is unremarkable on each margin but far from the joint cloud, and can flag a point that is extreme in x yet sits perfectly on the line and does no harm to the story. ## Diagnosing it **Plot first.** A scatterplot resolves this in one second and nothing else does. Quoting r without having seen the scatter is the underlying process failure; the outlier is just what it lets through. **Leave-one-out.** Recompute r n times, each time dropping one row, and record `max |r - r_without_i|`. This is cheap for any realistic n and gives a single number describing how much the headline depends on the single most influential row. A statistic that moves from 0.82 to 0.06 when one row leaves is not a finding, it is an anecdote about one row. **Report the sensitivity, don't hide it.** If a row is influential, publish r both ways and describe the row. Deleting the point and reporting the survivor is the failure mode reviewers look for. **Resample.** A bootstrap interval for r widens sharply when one point drives the estimate, because resamples that omit it look completely different from resamples that include it twice. A wide or bimodal interval is a fingerprint of the problem. ## Deciding what to do Influence is a diagnosis, not a verdict. Work through the row itself: - **Is it a data error?** Unit mix-ups, sentinel values such as -999 or a placeholder date, duplicated rows, a decimal shifted. Fix or drop it and say so. - **Is it a different population?** One wholesale account among retail customers, one instrument on a different scale, one day with a market halt. The honest response is usually to segment or to state the scope, not to keep a single mixed-population number. - **Is it real and rare?** Then the data genuinely cannot pin down the linear association well, and the right output is a range plus a description, not a confident single r. Sometimes the answer is to collect more data in that region. Transformations sometimes help: when a variable is strongly right-skewed and its extreme values are just the long tail rather than errors, analysing it on a log scale can compress the tail so no single row dominates -- but that changes what "linear" means and must be stated. ## Sample size The damage scales with 1/n. With n = 15, one row is nearly 7 percent of the evidence and can plausibly move r by a large amount; with n = 20,000 a single row of ordinary magnitude cannot, though a single row with an absurd value still can, because the contribution is unbounded. Correlations quoted from small samples deserve a leave-one-out check as a default habit rather than as a special investigation. ## In an interview Explain the mechanism first -- r is a sum of unbounded deviation products, so one far point can outweigh all the rest -- and only then list the diagnostics. Naming leave-one-out and "plot it" is table stakes; the senior signal is that you treat influence as something to report and investigate, not as licence to delete a row.

  • How would you quantify how much one observation is driving a reported r?
    Leave-one-out: recompute r n times, each with one row removed, and report the largest absolute change. That single number tells a reader how fragile the headline is. A bootstrap interval for r does the same job in aggregate -- it widens or turns bimodal when one row is doing the work, because resamples that exclude it are qualitatively different from resamples that include it.
  • Is it acceptable to delete the influential point and report the remaining correlation?
    Only with a stated reason that is about the row rather than about the result. A recorded unit error, a sentinel value or a row from a different population justifies removal, documented in the write-up. Removing a point because it is inconvenient is result-driven filtering. When the row is real and rare, publish r both ways and let the reader see how much rests on it.
  • Can screening each column separately for outliers catch these points?
    Not reliably. A pair can be unremarkable on both margins yet sit far from the joint cloud, and conversely a point extreme in x can lie exactly on the line the rest of the data follows and harm nothing. Influence on r is a joint property of the pair, so the check has to be joint too -- the scatterplot and leave-one-out both operate on pairs.
  • Why are small samples so much more exposed to this?
    Each row carries weight of roughly 1/n in the sum, so with n = 15 one row is close to 7 percent of the evidence before its extremity is even considered. As n grows, an ordinary row cannot move r much. The contribution is unbounded, though, so even a very large sample can be moved by one absurd value such as an unconverted unit or a sentinel code.

saying these in an interview costs you the question

  • Assumes one point out of thirty cannot matter
  • Quotes r without ever plotting the scatter
  • Deletes the influential row and reports only the survivor
  • Screens each column separately for extreme values
  • Treats an influential point as automatically a data error
  • Believes a large sample makes r immune to extreme values

context