skip to content

What does Anscombe's quartet show about trusting Pearson's r and other summary statistics?

level: middleimportance: must knowfreq 62%

answer

  1. four datasets, identical numbers
  2. eleven points each, r about 0.82
  3. same fitted line, four pictures
  4. curve, outlier, single leverage point
  5. summaries are not sufficient statistics

basics

~20 s

Anscombe's quartet is four eleven-point datasets sharing the same means, variances, correlation near 0.82 and fitted line to two decimals, yet their scatterplots look completely different. Summary statistics do not identify a dataset - plot before concluding.

solid answer

~50 s

The quartet is four small datasets constructed so that the mean and variance of x, the mean and variance of y, the correlation near 0.82 and the least-squares line all agree to two decimal places. Plotted, they are nothing alike: the first is an ordinary noisy straight-line cloud, the second is a clean curve that a straight line badly misrepresents, the third is a perfect line with a single point pulled off it, and the fourth has every x at one value except one far-away point that alone creates the slope and the correlation. The lesson is that summary statistics compress and the compression is lossy: identical numbers, four different stories, three of which make the fitted line meaningless. The modern descendant is the Datasaurus set, where wildly different pictures — including a dinosaur — share summary statistics. The practical rule is to look at the scatterplot before quoting r.

go deeper

for a junior

Be ready to say what the quartet is — four datasets with matching summary numbers and completely different scatterplots — and to conclude that you should plot data before summarising it.

for a middle

Explain what each panel breaks: a curve defeats the linearity that r assumes, a lone outlier moves the fit, and a single distant x-value manufactures a slope from nothing.

for a senior

Show how you would catch these without plotting every pair: residual structure, per-observation influence, and whether the predictor genuinely varies where the conclusion is applied.

for a principal

Own the standard for what a report must show alongside a coefficient, so that team conclusions are never defensible only because nobody looked at the shape of the data.

## The construction Francis Anscombe published four datasets in 1973, each with eleven `(x, y)` pairs, deliberately built so that a long list of summary statistics is shared to two decimal places: - the mean of `x` (9) and the variance of `x` (11) - the mean of `y` (7.50) and the variance of `y` (about 4.13) - the correlation between `x` and `y` (about 0.816) - the least-squares line, roughly `y = 3.00 + 0.500 * x` - consequently `r^2`, about 0.67 A report containing only those numbers would describe all four identically. Every automated summary would call them the same dataset. ## The four pictures **Dataset I** is what the numbers suggest: a straight-line relationship with ordinary scatter around it. Here the summaries are honest, and the fitted line is a fair description. **Dataset II** is a smooth concave curve with essentially no noise. The relationship is nearly deterministic, but it is not a straight line, so the fitted line runs through the middle of a bend — over-predicting at the ends and under-predicting in the centre in a completely systematic way. The correlation of 0.816 understates how tightly the two variables are tied and misdescribes the shape entirely. **Dataset III** is a perfect straight line through ten points, with one point displaced far off it. The outlier drags the fitted line away from the line that ten of the eleven points define, and it drags the correlation down from 1. Here the right description is 'a deterministic law plus one anomalous observation', which no summary statistic reports. **Dataset IV** is the most alarming. Ten points share one identical `x` value; a single eleventh point sits far to the right. There is no evidence at all about how `y` responds to `x` — the data contains exactly one distinct x-value pair to work with — yet a slope, a correlation and an `r^2` are dutifully computed. Delete the one distant point and the slope is undefined. The entire relationship is one observation. ## Why this is more than a curiosity The interview point is not the anecdote; it is what each panel breaks: - Panel II breaks the **linearity assumption**. Pearson's `r` and least squares both measure straight-line structure; a curve is invisible to them as a curve and merely shows up as a middling coefficient. - Panel III breaks **robustness**. A tiny number of extreme points can move an average of products a long way, because deviations are not bounded. - Panel IV breaks **support**. A model can only speak about the region where you have data spread out; leverage from one distant point produces a confident-looking estimate with nothing behind it. All three failures are silent. Nothing in the summary output flags them. That is the actual claim: summary statistics are not sufficient statistics for an arbitrary dataset, so equal numbers do not mean equal data. ## The modern version The Datasaurus Dozen extends the demonstration: starting from a scatterplot shaped like a dinosaur, a set of very different shapes — stars, lines, clusters, an X — is produced while holding the means, standard deviations and correlation fixed to two decimals throughout. It makes the same point at larger `n`, which is worth knowing because people sometimes assume the quartet only works because eleven points is a tiny sample. It does not depend on small `n`. ## What to do about it in practice The habit is cheap: for any pair you are about to summarise with a single coefficient, look at the scatterplot first. At scale, where plotting every pair is impossible, the practical substitutes are automatic: check residuals against the fitted values for structure rather than trusting `r^2`; look at the influence of individual observations rather than assuming none dominates; check that the predictor actually varies across the range where the conclusion will be applied. None of these require a plot per pair, but all of them are prompted by understanding what the quartet exposes. A candidate who can name the four panels and say which modelling assumption each one violates has demonstrated more than recall of a famous picture — they have shown they know what a correlation coefficient is silent about.

  • Which of the four panels is the most dangerous in real work, and why?
    The fourth. In the curve and outlier cases the data still carries information about the relationship, just not the information the line reports. In the fourth panel the predictor is essentially constant apart from one distant observation, so the slope and the correlation rest entirely on a single point. It looks like an estimate and is really an artefact of one row.
  • Does the demonstration only work because each dataset has eleven points?
    No. Small n makes the quartet easy to construct by hand, but the same trick scales: the Datasaurus set holds means, standard deviations and correlation fixed across dramatically different shapes at much larger sample sizes. Summary statistics compress away shape at any n, so more data does not rescue a summary-only reading.
  • You cannot plot ten thousand variable pairs. What do you check instead?
    Automatable proxies for what the plot would show: residual patterns against fitted values to catch systematic curvature, per-observation influence to catch a single dominating point, and the spread of the predictor to catch estimates resting on a narrow or degenerate range. These are the numeric versions of the three failures the quartet exposes.

Four very different books can have the same page count, word count and average sentence length. Those numbers tell you nothing about which one is a novel and which is a phone directory.

saying these in an interview costs you the question

  • Thinks identical summary statistics imply similar data
  • Believes r near 0.8 always means a good straight-line fit
  • Claims the quartet is an artefact of eleven points
  • Reports r without ever seeing the scatterplot
  • Treats a high r-squared as proof the model fits

context