skip to content

How does a boxplot's 1.5 x IQR rule decide that a data point is an outlier?

level: middleimportance: must knowfreq 66%

answer

  1. quartiles, not the mean and SD
  2. a multiple of the box width
  3. whiskers stop at real data, not at the cutoff
  4. about 0.7 percent flagged under a normal curve

basics

~20 s

It takes the first and third quartiles Q1 and Q3, sets IQR = Q3 - Q1, and marks any point below Q1 - 1.5 x IQR or above Q3 + 1.5 x IQR. Those cutoffs are called the fences.

solid answer

~50 s

You compute the quartiles, take `IQR = Q3 - Q1`, and place fences at `Q1 - 1.5 * IQR` and `Q3 + 1.5 * IQR`. Anything outside the fences is plotted as an individual point. A detail candidates often miss: the whiskers do not extend to the fences — they stop at the most extreme actual observation still inside them, so whisker length varies with the data. The 1.5 multiplier is a convention chosen by Tukey, not a significance level: under a normal distribution the fences sit at about 2.7 standard deviations, so roughly 0.7% of clean points fall outside, which is about seven flagged points in a clean sample of a thousand. Tukey also defined an outer fence at 3 x IQR for far-out points. The rule is a display convention that says unusual, and it never says wrong.

go deeper

for a junior

Know that a boxplot's box spans the quartiles and that points drawn separately beyond the whiskers are the ones the rule has flagged as unusual.

for a middle

State the fences exactly as Q1 minus 1.5 times the IQR and Q3 plus 1.5 times the IQR, and know that whiskers stop at the last real point inside them.

for a senior

Show judgment about when the rule misleads: asymmetric columns, pooled groups and tiny samples all produce flag counts that say more about shape than about data quality.

for a principal

Own the argument that 1.5 is a convention with no error-rate guarantee, and decide when a team should tune or abandon a fixed threshold rather than apply it everywhere.

## The construction A boxplot is built from five numbers. Sort the data, take the first quartile `Q1` (the value below which a quarter of the data falls), the median, and the third quartile `Q3`. The box spans `Q1` to `Q3` and is crossed by the median. The interquartile range is `IQR = Q3 - Q1`, the width of that box. The fences are then - lower fence: `Q1 - 1.5 * IQR` - upper fence: `Q3 + 1.5 * IQR` Every observation strictly outside a fence is drawn as its own marker. Everything inside is summarised by the whiskers. ## The detail people get wrong The whiskers are not drawn at the fences. Each whisker extends from the box to the most extreme observation that is still inside the fence on that side. If your largest non-flagged value sits well short of the upper fence, the upper whisker is short. This is why two boxplots with identical boxes can have very different whisker lengths, and why you cannot read the fence position off the plot directly — you can only read where real data stopped. ## Why 1.5 The multiplier is a convention introduced by John Tukey, chosen so that the rule flags few points on well-behaved data while remaining simple enough to apply by hand. It has a clean interpretation under a normal distribution: quartiles sit at about `0.6745` standard deviations either side of the centre, so `IQR` is about `1.349` standard deviations, and the fences land at about `0.6745 + 2.0235 = 2.698` standard deviations from the centre. The probability of a normal observation falling beyond that is roughly 0.7% in total across both tails. In a clean normal sample of 1,000 points you should therefore expect about seven flagged points and be entirely unsurprised by them. Tukey also defined outer fences at `Q1 - 3 * IQR` and `Q3 + 3 * IQR`, sometimes drawn with a different marker; points beyond those he called far out. Nothing about either multiplier is a hypothesis test — there is no null distribution, no p-value and no error-rate guarantee attached to the rule. ## Why it is robust The fences depend only on `Q1` and `Q3`. Both are quantiles, so both are functions of rank rather than magnitude, and both can absorb a large amount of contamination before they move. That is the rule's key advantage over a mean-and-standard-deviation cutoff: the extreme values you are hunting for do not participate in setting their own threshold. A cutoff built from the mean and standard deviation is contaminated by exactly the points it is meant to catch. ## Where it misleads: asymmetric data The rule is symmetric about the box, but many real distributions are not. Response times, incomes, session lengths and claim sizes have a long right tail by nature. On such data the upper fence sits close to the bulk in relative terms while the tail extends far past it, so the rule flags a fistful of points — sometimes dozens — none of which are errors. A clean, right-tailed sample can easily produce ten or twenty flagged points, and the correct reading is that the distribution is asymmetric, not that the data is dirty. The practical consequences: - The count of flagged points is a statement about distribution shape at least as much as about data quality. - Applying the rule inside groups can be very different from applying it to a pooled column, because pooling several differently-centred groups widens `IQR` and hides genuine within-group extremes. - On tiny samples the quartiles themselves are unstable, so the fences jump around; with fewer than about ten points the rule tells you very little. ## What to do with a flagged point The flag is a prompt to look, not a licence to delete. A flagged point deserves a provenance check — where did the value come from, is it physically possible, does an independent source corroborate it — before anyone decides it is contamination. Deleting on the strength of the fence rule alone systematically shaves real tail mass off asymmetric data, and tail behaviour is often exactly what the analysis was about. ## What interviewers listen for The formula stated correctly with quartiles rather than the mean, the whisker detail, an awareness that 1.5 is a convention with roughly 0.7% coverage under normality, and the judgment to say that skewed data will trigger the rule without anything being wrong.

  • Why is a fence built from quartiles safer than a cutoff built from the mean and standard deviation?
    Because the quartiles barely move when a few extreme values are present, so the threshold is not set by the points it is trying to catch. A mean-and-standard-deviation cutoff is computed from the contaminated sample: the extreme value pulls the centre toward itself and inflates the spread, widening the very threshold meant to flag it. Quartiles depend on rank, so contamination in the tails leaves them essentially where they were.
  • A clean, right-tailed column of response times flags fifteen points beyond the upper fence. What do you conclude?
    That the distribution is asymmetric, not that fifteen measurements are wrong. The rule is symmetric around the box, so a long right tail naturally produces points past `Q3 + 1.5 * IQR`. The right response is to look at the distribution as a whole and at where those values came from, and to resist any rule that would delete genuine tail mass simply because a display convention drew them separately.
  • What is the outer fence at 3 x IQR for?
    It marks far-out points: values beyond `Q1 - 3 * IQR` or `Q3 + 3 * IQR`. Tukey defined it as a second, stricter tier so a reader can distinguish mildly unusual points from truly distant ones, and some plotting conventions draw the two tiers with different markers. Like the 1.5 multiplier it is a convention with no error-rate guarantee attached, just a coarser filter for the points most worth investigating first.

saying these in an interview costs you the question

  • States the fences using the mean and standard deviation instead of quartiles
  • Says the whiskers are drawn at the fence positions
  • Treats 1.5 x IQR as a statistical significance threshold
  • Assumes every flagged point is an error to delete
  • Applies the rule to tiny samples where quartiles are unstable

context