skip to content

On a normal Q-Q plot of a raw sample, what do an S-shape and a single bent tail mean?

level: middleimportance: should knowfreq 48%

answer

  1. sorted data against expected positions
  2. read the ends, not the middle
  3. both ends bending the same way
  4. opposite bends at the two ends
  5. small samples wobble even when normal

basics

~20 s

An S-shape means both tails are heavier than a normal's: the lowest points sit below the reference line and the highest above it. A curve where both ends bend the same way, one tail far more than the other, means skew.

solid answer

~50 s

A normal Q-Q plot sorts the sample and plots each value against the value a normal distribution would place at that position - sample quantiles on the vertical axis, theoretical quantiles on the horizontal. Points hugging the straight reference line mean the shapes agree. An S-shaped pattern, with the lowest points below the line and the highest points above it, says both tails run further out than a normal's: heavy tails. If instead both ends deviate in the same direction - the upper tail arcing well above the line while the compressed lower end also sits above it - the sample is right-skewed, and the mirror image is left skew. I read the ends, not the middle: wobble in the centre is ordinary sampling noise, and at small n even a perfectly normal sample wanders around the line.

go deeper

for a junior

Be ready to say what the two axes are - sample quantiles against the values a normal would predict - and that points near the straight line mean the shapes agree.

for a middle

Explain the construction from sorted values and plotting positions, and map each canonical curve to its cause: opposite bends for heavy tails, a same-direction arc for skew, a staircase for rounding.

for a senior

Connect the shape to a consequence. Say which procedures a heavy tail actually threatens, when moderate skew is absorbed by sample size, and when you would judge the departure by resampling instead of by eye.

for a principal

Set the norm for how shape evidence is reported and acted on, so teams do not either ignore plots entirely or block decisions on cosmetic bends that no downstream number is sensitive to.

## How the plot is built A quantile is a value below which a given proportion of a distribution falls: the 0.9 quantile is the point with 90% of the mass beneath it. A normal Q-Q plot pairs the quantiles of your sample with the quantiles a normal distribution would have at the same positions. Concretely: sort the n observations; the i-th smallest is treated as the empirical quantile at roughly the (i - 0.5)/n position; look up the standard normal value at that same position; plot the pair. By convention the sample quantiles go on the vertical axis and the theoretical ones on the horizontal. A reference line is drawn through the points - commonly through the first and third quartile pair - so that a sample whose shape matches a normal falls along it regardless of its mean and standard deviation. The plot is about shape, not about location or scale. ## Reading the four canonical departures **Heavy tails.** If the sample's extremes are more extreme than a normal's, the smallest points are more negative than predicted (below the line) and the largest are more positive than predicted (above the line). The two ends bend in opposite vertical directions, producing the classic S-shape. This is what a t-like distribution with few degrees of freedom, or a contaminated sample, looks like. **Light tails.** The reverse S: the low end sits above the line and the high end below it, because the sample's extremes are milder than a normal's. Bounded or truncated measurements often look like this. **Right skew.** The long tail is on the high side, so the largest values overshoot dramatically and the points at the right end climb well above the line. On the low side the distribution is compressed relative to a normal, so those points also sit above the line. Both ends deviate the same way and the picture reads as a single upward-curving arc, dominated by the right tail. Income, session duration, revenue per customer and waiting times routinely look like this. **Left skew.** The mirror image: both ends below the line, with the lower tail dropping furthest, giving a downward-curving arc. ## What is not a departure Order statistics are noisy. With n = 25 a perfectly normal sample produces visible wandering, worst at the two extreme points where a single draw fixes the position. The habit that separates a competent reader from a nervous one is to look for a systematic, monotone bend rather than local wiggle, and, when uncertain, to compare against a handful of simulated normal samples of the same size to calibrate the eye. Discrete or rounded data produce a staircase of horizontal runs; that is a property of the measurement, not evidence about the underlying shape. A handful of points detached at one end is an outlier signal, and is often more actionable than the overall shape. ## Why plot rather than test A plot answers the question you actually have. It shows the direction of the departure (skew versus tail weight), its magnitude, and whether it is driven by a few points or by the bulk. A formal normality test compresses all of that into a single verdict about a hypothesis nobody believes exactly, and the verdict is driven heavily by sample size. Knowing that the departure is one long right tail suggests a different response - a transformation, a different summary, more caution about the mean - than knowing that both tails are heavy, which is mostly a warning about outlier influence and power. ## Turning the reading into a decision The follow-up question is always: does this shape threaten the specific procedure? For a comparison of means at a healthy sample size, moderate skew is largely absorbed by the central limit theorem and the plot's message is mainly about outliers. For a small sample, the raw shape matters directly, because the central limit theorem has not had room to work. For anything that depends on the tails themselves - a high percentile, a prediction interval, a threshold-crossing probability - the tail behaviour the plot shows is exactly the thing being estimated, and a heavy-tail S-shape is a serious warning that a normal-based answer will understate the extremes. Say all of that in an interview and you have shown the skill being probed: not memorising which curve is which, but connecting a shape to a consequence.

  • What does a normal Q-Q plot look like for left-skewed data?
    It bends the opposite way from right skew: both ends fall below the reference line, with the lower tail dropping furthest because the long tail is on the low side. The upper end is compressed relative to a normal, so those points also sit under the line. Read as a curve, it is a downward arc rather than an upward one.
  • At n = 25 the middle points wander off the line. Is that evidence of non-normality?
    Rarely. Order statistics are noisy, and at n = 25 even a perfectly normal sample shows visible wobble, worst at the extreme points where one draw sets the position. I look for a systematic, monotone bend that persists across the plot rather than local wiggle, and if I need reassurance I compare against a few simulated normal samples of the same size.
  • Why prefer a Q-Q plot over a formal normality test?
    The plot shows the direction and magnitude of the departure, which is what decides whether it matters. A test returns only a verdict on a hypothesis nobody believes exactly, and that verdict tracks sample size more strongly than it tracks any practically relevant deviation. Seeing that the problem is one long right tail tells me what to do next; a p-value does not.

saying these in an interview costs you the question

  • Reads the middle of the plot and ignores the tails
  • Calls any deviation from the line non-normal, whatever the sample size
  • Confuses a skew arc with a heavy-tail S-shape
  • Thinks the reference line is fitted in order to prove normality
  • Treats a staircase caused by rounded data as evidence about shape

context