How can a single extreme point push Pearson's r from near zero to 0.8?
answer
- the statistic is a sum over rows
- each term multiplies two deviations
- far from both means, both deviations large
- one term outweighs all the others
- drop each row in turn and recompute
basics
~20 sPearson's r sums cross-products of deviations, so one point far from both means can contribute a term that outweighs all the others and dominates the statistic. Plot the scatter and recompute r with each row removed to catch it.
solid answer
~50 sr sums the products `(xi - xbar) * (yi - ybar)` and divides by the two standard deviations. A point far from both means contributes a term proportional to the product of two large deviations, so with 30 unstructured points and one point far out along a line, that single term can outweigh the other 30 combined and drag r from near 0 to 0.8 or beyond. The lever works in reverse too: one point off the axis of an otherwise tight relationship can collapse a high r toward 0. Diagnosis is mechanical. Look at the scatterplot before quoting r, then recompute r with each row dropped in turn and record the largest swing. If one row moves r materially, report r with and without it and investigate whether that row is a data error or a genuine rare case.
go deeper
Know that Pearson's r is fragile to extreme values and that a scatterplot should be looked at before any correlation is quoted. Recall that one point far from the rest can raise or lower r substantially.
Explain the mechanism: r sums products of two deviations, so a point far from both means contributes a term that can outweigh every other. Describe leave-one-out recomputation as the concrete check.
Demonstrate the full loop -- detect, investigate the row, decide, and disclose. Interviewers want to hear that you report r with and without the point rather than quietly dropping it, and that you distinguish a data error from a genuine rare case.
Own the standard: which association numbers ship with a sensitivity check, what happens when a headline metric turns out to rest on one row, and how the team avoids a culture of quietly filtering data until the correlation looks right.
## Why one point has so much leverage Pearson's r for n paired observations is `r = sum_i (xi - xbar) * (yi - ybar) / ((n - 1) * sx * sy)` The numerator is a sum of **products of deviations**. Two things make that sum easy for one row to hijack: 1. **The contribution is a product, not a distance.** A point ten deviations from the mean in x and ten in y contributes roughly a hundred times the term of a typical point sitting one deviation out in each direction. 2. **The sum is unweighted and unbounded.** Nothing caps how much one term can add. Push a single point far enough along any straight line and r converges toward plus or minus 1, no matter what the other n - 1 points look like. With 30 points forming a formless blob (r near 0) and one added point far away along a line through the blob, that point's cross-product can exceed the total from the other 30, and the reported r lands at 0.8 or higher. The same mechanism runs backwards: take 30 points lying tightly along a line, add one point far away perpendicular to it, and r can fall toward 0. Identical data, one row of difference, opposite headline. ## Two different kinds of extreme point It helps to separate them: - A point extreme in **both** variables and consistent with a straight line through the data inflates r. This is the classic high-leverage point. - A point extreme in **one** variable but ordinary in the other, or extreme in a direction the rest of the data does not follow, deflates r. The distinction matters when you go looking: scanning for outliers one column at a time can miss a pair that is unremarkable on each margin but far from the joint cloud, and can flag a point that is extreme in x yet sits perfectly on the line and does no harm to the story. ## Diagnosing it **Plot first.** A scatterplot resolves this in one second and nothing else does. Quoting r without having seen the scatter is the underlying process failure; the outlier is just what it lets through. **Leave-one-out.** Recompute r n times, each time dropping one row, and record `max |r - r_without_i|`. This is cheap for any realistic n and gives a single number describing how much the headline depends on the single most influential row. A statistic that moves from 0.82 to 0.06 when one row leaves is not a finding, it is an anecdote about one row. **Report the sensitivity, don't hide it.** If a row is influential, publish r both ways and describe the row. Deleting the point and reporting the survivor is the failure mode reviewers look for. **Resample.** A bootstrap interval for r widens sharply when one point drives the estimate, because resamples that omit it look completely different from resamples that include it twice. A wide or bimodal interval is a fingerprint of the problem. ## Deciding what to do Influence is a diagnosis, not a verdict. Work through the row itself: - **Is it a data error?** Unit mix-ups, sentinel values such as -999 or a placeholder date, duplicated rows, a decimal shifted. Fix or drop it and say so. - **Is it a different population?** One wholesale account among retail customers, one instrument on a different scale, one day with a market halt. The honest response is usually to segment or to state the scope, not to keep a single mixed-population number. - **Is it real and rare?** Then the data genuinely cannot pin down the linear association well, and the right output is a range plus a description, not a confident single r. Sometimes the answer is to collect more data in that region. Transformations sometimes help: when a variable is strongly right-skewed and its extreme values are just the long tail rather than errors, analysing it on a log scale can compress the tail so no single row dominates -- but that changes what "linear" means and must be stated. ## Sample size The damage scales with 1/n. With n = 15, one row is nearly 7 percent of the evidence and can plausibly move r by a large amount; with n = 20,000 a single row of ordinary magnitude cannot, though a single row with an absurd value still can, because the contribution is unbounded. Correlations quoted from small samples deserve a leave-one-out check as a default habit rather than as a special investigation. ## In an interview Explain the mechanism first -- r is a sum of unbounded deviation products, so one far point can outweigh all the rest -- and only then list the diagnostics. Naming leave-one-out and "plot it" is table stakes; the senior signal is that you treat influence as something to report and investigate, not as licence to delete a row.
- How would you quantify how much one observation is driving a reported r?Leave-one-out: recompute r n times, each with one row removed, and report the largest absolute change. That single number tells a reader how fragile the headline is. A bootstrap interval for r does the same job in aggregate -- it widens or turns bimodal when one row is doing the work, because resamples that exclude it are qualitatively different from resamples that include it.
- Is it acceptable to delete the influential point and report the remaining correlation?Only with a stated reason that is about the row rather than about the result. A recorded unit error, a sentinel value or a row from a different population justifies removal, documented in the write-up. Removing a point because it is inconvenient is result-driven filtering. When the row is real and rare, publish r both ways and let the reader see how much rests on it.
- Can screening each column separately for outliers catch these points?Not reliably. A pair can be unremarkable on both margins yet sit far from the joint cloud, and conversely a point extreme in x can lie exactly on the line the rest of the data follows and harm nothing. Influence on r is a joint property of the pair, so the check has to be joint too -- the scatterplot and leave-one-out both operate on pairs.
- Why are small samples so much more exposed to this?Each row carries weight of roughly 1/n in the sum, so with n = 15 one row is close to 7 percent of the evidence before its extremity is even considered. As n grows, an ordinary row cannot move r much. The contribution is unbounded, though, so even a very large sample can be moved by one absurd value such as an unconverted unit or a sentinel code.
saying these in an interview costs you the question
- Assumes one point out of thirty cannot matter
- Quotes r without ever plotting the scatter
- Deletes the influential row and reports only the survivor
- Screens each column separately for extreme values
- Treats an influential point as automatically a data error
- Believes a large sample makes r immune to extreme values