In a normal Q-Q plot of regression residuals, which pattern signals heavy tails rather than right skew?
answer
- state the axis convention first
- ends in opposite directions versus the same direction
- tail weight versus asymmetry
- convex curve means a long upper tail
- extreme points are the noisiest
basics
~20 sHeavy tails bend both ends away from the straight line in opposite directions: the lowest points fall below it and the highest rise above it. Right skew bends the whole plot upward, with both ends above the line and the largest gap at the top.
solid answer
~50 sTake the usual convention: ordered residuals on the vertical axis, quantiles of a standard normal on the horizontal axis. If the residuals really are normal the points follow a straight line. Heavy tails mean the extreme residuals are more extreme than a normal distribution predicts at both ends, so the low end drops below the line and the high end rises above it, giving a stretched S. Right skew means the upper tail is long while the lower tail is compressed, so the low points sit above the line too, and the whole curve is convex, bending upward with its biggest departure at the top. Left skew is the mirror image, concave. Always state your axis convention first, because flipping the axes reverses the reading, and treat wobble at the very ends of a small sample as noise rather than evidence.
go deeper
Know what the plot shows: sorted residuals against the quantiles a normal distribution would produce, and points on a straight line means consistent with normal. Being able to say curvature means non-normal is enough here.
Explain the construction with quantiles, state your axis convention, and read the two canonical departures correctly: ends flaring in opposite directions means heavy tails, both ends bending the same way means skew.
Demonstrate calibration judgment: know how much wobble is expected at a given sample size, resist over-diagnosing from a couple of extreme points, and connect a skewed or heavy-tailed picture to a modelling cause rather than a cosmetic fix.
Decide what evidence of non-normality should change. Set the expectation that departures are judged by size and consequence for the intended use, and steer teams away from ritual normality checking that never alters a decision.
## How a Q-Q plot is built Sort the `n` residuals from smallest to largest. The `i`-th sorted residual is the empirical quantile at roughly the fraction `(i - 0.5) / n` of the data. Look up the quantile of a standard normal distribution at that same fraction. Plot the pair. Do this for every observation and you have a normal Q-Q plot: sample quantiles on the vertical axis, theoretical normal quantiles on the horizontal. If the residuals come from a normal distribution, sample and theoretical quantiles move together and the points lie on a straight line. The line's slope reflects the residual standard deviation and its intercept the mean, so the *straightness* is what matters, not whether the line happens to be the 45-degree diagonal. **State the convention before you interpret.** Some tools put theoretical quantiles on the vertical axis; every up/down reading below then flips. Saying which axis is which is part of a correct answer. ## Heavy tails A heavy-tailed distribution is symmetric-ish but produces extreme values more often than a normal with the same spread. In quantile terms, its low quantiles are more negative and its high quantiles are more positive than the normal's. On the plot: the points at the far left lie **below** the line, the points at the far right lie **above** it. The middle of the plot tracks the line reasonably well. The result is a shallow S rotated so both ends flare away from the line in opposite vertical directions. Practical reading: a handful of observations are much further from the fit than the bulk. Least squares is sensitive to those, so a small number of rows carry disproportionate weight in the estimate, and any interval that depends on the shape of the error distribution rather than on an average is affected. ## Right skew A right-skewed distribution has a long upper tail and a short, bunched lower tail. Its high quantiles are more extreme than the normal's, but its low quantiles are **less** extreme — closer to the centre. On the plot: the far-right points lie above the line, and the far-left points also lie above the line, because they are not as negative as a normal would predict. The whole curve is convex — it bends upward — and the biggest departure is at the top. It looks like one bend, not two flares. That single distinguishing test is worth memorising: **opposite directions at the two ends means tail weight; the same direction at both ends means asymmetry.** ## The other two shapes - **Light tails** (a distribution more compact than normal, sometimes from a bounded outcome): the left end sits above the line and the right end below it — the S in the opposite orientation to heavy tails. This is the least worrying pattern for inference; intervals built assuming normality are conservative. - **Left skew**: the mirror of right skew, concave, both ends below the line, biggest departure at the bottom. A discrete or heavily rounded outcome produces yet another signature: horizontal steps, because many residuals share exactly the same value. ## How much departure is real The extreme order statistics are the noisiest points on the plot. Even with genuinely normal residuals, the largest of 30 observations can sit visibly off the line; with `n = 20` you should expect the ends to wander. This is why so many people over-diagnose non-normality: they read the two or three most eye-catching points instead of the systematic shape. Two disciplines help. First, judge the *shape*, not individual points: is the deviation systematic and in a consistent direction over many points, or is it one point at the very end? Second, calibrate by simulation: generate several samples of the same size from a normal distribution, draw their Q-Q plots, and see how much wobble is normal at that `n`. Some tools draw this directly as a band around the reference line. At the other end, with tens of thousands of observations the plot will show real but tiny departures with great clarity. Then the question stops being 'is it exactly normal' — nothing is — and becomes 'is the departure big enough to matter for what I am using the model for'. ## What to do with the finding The Q-Q plot is a description, not a verdict. Right skew in the residuals often means the response should be modelled on a different scale, or that a predictor explaining the long tail is missing. Heavy tails often mean a mixture of populations that the model treats as one, or genuine rare large deviations that ordinary least squares handles poorly because it squares them. One caution that catches people out: the plot is about the **residuals**, not the response. A strongly skewed response variable is perfectly compatible with well-behaved residuals, because the predictors may explain exactly that skew. Checking the marginal distribution of the outcome instead of the residuals is a classic mistake.
- Why do the extreme points wander off the line even when the residuals really are normal?Because the largest and smallest order statistics have far more sampling variability than the middle ones. With thirty observations, the top point sitting noticeably off the line is unremarkable. Judge the systematic shape across many points, and if you need calibration, compare with Q-Q plots of simulated normal samples of the same size.
- What does a Q-Q plot of light-tailed residuals look like, and does it matter?The S runs the other way: the low points sit above the line and the high points below it, because the extremes are less extreme than a normal predicts. It rarely matters for inference — intervals built on a normal assumption end up conservative rather than anti-conservative.
- The residuals show clear right skew. What is usually the underlying cause?Most often the response is naturally multiplicative or bounded below, so its errors are asymmetric on the raw scale, or a predictor that explains the long upper tail is missing. Modelling the response on a log scale, or adding the missing structure, typically straightens the plot better than any post-hoc adjustment.
- Does a skewed response variable by itself mean the residuals will be non-normal?No. The assumption concerns the errors, not the marginal distribution of the outcome. If the predictors explain the skew, the residuals can be perfectly symmetric. Checking the histogram of the response instead of the residuals is a common and consequential mix-up.
Think of the reference line as a ruler and the points as a bent rod: heavy tails bend the two ends the opposite way, like an S, while skew bends the whole rod one way, like a banana.
saying these in an interview costs you the question
- Reads a single stray point at the end as non-normality
- Checks the distribution of the response instead of the residuals
- Cannot say which axis holds the theoretical quantiles
- Calls any bend heavy tails without checking both ends
- Assumes any departure from the line invalidates the model