What is the five-number summary of a dataset, and what does it tell you that the mean hides?
answer
- Five positions, not five averages
- Sorted data, split into quarters
- Median plus two quartiles plus both ends
- Mean below median means a low tail
basics
~20 sThe five-number summary is the minimum, first quartile, median, third quartile and maximum. Together they show where the bulk of the data sits and how far each tail reaches, which a single average cannot express.
solid answer
~40 sThe five-number summary reports five order-based values from the sorted data: the minimum, the first quartile Q1 (the 25th percentile), the median (50th), the third quartile Q3 (75th) and the maximum. A percentile is simply the value below which that share of the sorted observations falls, so half the data sits below the median and the middle half sits between Q1 and Q3. Take a class of exam scores with mean 72 and summary 20 / 65 / 78 / 85 / 98: the mean alone suggests a mediocre class, but the summary shows the typical student scored 78 and the mean is dragged down by a low tail reaching 20. Because the four inner values are positions in the sorted list, changing the worst score from 20 to 2 moves only the minimum.
go deeper
Be ready to name all five values in order and compute them on a short sorted list, including the even-count median rule. Knowing that the median is the 50th percentile is the screening bar here.
Explain why a gap between the mean and the median signals an asymmetric tail, and why moving one extreme value changes the mean but not the median. Show that equal counts per quarter does not mean equal widths.
Demonstrate judgment about which number to put in front of stakeholders. Say when a median plus a tail value is the honest summary of a business metric, and when a single average would actively mislead a decision.
Own the reporting convention: decide what your organisation's dashboards default to, and whether teams are allowed to publish a bare mean at all. The tradeoff is comparability across teams against faithfulness to each metric's shape.
## What a quantile is Sort the observations from smallest to largest. The **p-th percentile** is the value that separates the lowest p percent of the sorted data from the rest: about p percent of observations are at or below it, and about (100 - p) percent are at or above it. Quantiles are the same idea on a 0-to-1 scale, so the 0.9 quantile and the 90th percentile are the same number. Because they are defined by *position in the sorted list* rather than by arithmetic on the values, quantiles describe where data sits rather than averaging it. Three percentiles get special names. The **median** is the 50th percentile, the value with half the data below it. The **first quartile Q1** is the 25th percentile and the **third quartile Q3** is the 75th; the middle half of the data lies between them. ## The five-number summary The five-number summary is the compact profile min, Q1, median, Q3, max. Five numbers, read left to right, sketch the whole shape: - **min to Q1** covers the lowest quarter of the data. - **Q1 to median** covers the second quarter. - **median to Q3** covers the third quarter. - **Q3 to max** covers the top quarter. Each interval holds the same *count* of observations (a quarter of them) but generally spans a different *width*. Comparing those widths is how you read shape from five numbers. If the top quarter spans a much wider range than the bottom quarter, the data stretch out to the right; if the reverse, they stretch to the left. ## A worked example A class of 100 exam scores has mean 72 and five-number summary 20 / 65 / 78 / 85 / 98. The mean of 72 on its own reads as a class that did adequately. The summary tells a sharper story. The median is 78, so half the class scored 78 or better. Twenty-five students scored 85 or better, and twenty-five scored 65 or worse, with the bottom of the range reaching 20. The bottom quarter spans 45 points (20 to 65) while the top quarter spans only 13 (85 to 98). That asymmetry is why the mean, at 72, sits six points below the median: a handful of very low scores pull the arithmetic average down, while the median simply counts positions and does not care how low the low scores go. The practical reading is that this is not a uniformly mediocre class; it is a mostly-strong class with a small group of students who are badly behind. Those are two completely different interventions, and only the summary distinguishes them. ## Why the mean alone is not enough The mean is a single number computed from all the values, so every observation influences it in proportion to its size. One extreme value can move the mean noticeably. The median, Q1 and Q3 are positions, so replacing the worst score of 20 with 2 changes only the minimum and leaves the other four numbers untouched. That is the core contrast an interviewer is testing: the mean answers 'what is the balance point of the values', the summary answers 'where are the observations actually located'. The mean also cannot express asymmetry or reach. Two datasets with identical means can have completely different summaries, and the summary is what tells you whether a typical case and a worst case are close together or far apart. ## What the five-number summary still cannot show It is five numbers, so it compresses hard. It cannot reveal **multimodality**: a dataset whose values cluster around two separate humps can produce exactly the same five numbers as one smooth mound, because both can put a quarter of their mass in each of the four intervals. It also says nothing about how the data are distributed *inside* each quarter. For that you need a fuller view of the whole distribution, such as a histogram or an empirical CDF, which plots the cumulative share of observations at or below each value and lets you read any percentile straight off the axis. ## Reporting practice When you report a summary statistic to a non-technical audience, quoting the median alongside the mean and at least one tail value is a cheap way to avoid the classic misread. If the mean and the median differ meaningfully, that gap is itself the headline, and the five-number summary is the shortest honest way to explain it.
- How do you compute the median of an even number of observations?Sort the values and take the two middle ones, then average them. With 10 observations the median is the average of the 5th and 6th values. With an odd count there is a single middle observation and no averaging is needed. Note that this makes the median of an even-sized sample a value that need not appear in the data.
- Two datasets have the same five-number summary. Can they still look different?Yes. The summary fixes only five positions; the data inside each quarter can be arranged any way at all. One dataset could pile its observations near the median while another splits into two separate clusters, and both can produce identical quartiles. Detecting that difference needs a histogram or an empirical CDF, not five numbers.
- When would you report the mean rather than the median?When you need a total or a rate that must add up: revenue per user times user count gives total revenue only with the mean, not the median. The mean is also the right choice when the distribution is roughly symmetric and you want every observation to contribute. For skewed data reported as a typical case, prefer the median.
The mean is the balance point of a seesaw; the five-number summary is a photo of where everyone is actually sitting on it.
saying these in an interview costs you the question
- Calls the median the average of the smallest and largest values
- Says the median is always close to the mean
- Thinks each quarter of the data spans an equal range of values
- Claims the five-number summary can reveal two separate humps
- Confuses the first quartile with the first 25 observations