What does the sign of sample skewness tell you about the shape of a distribution?
answer
- think about which side stretches out
- third standardised moment, deviations cubed
- cubing keeps the sign of a deviation
- the number names where the tail points
- long right tail gives a positive value
basics
~20 sPositive sample skewness means the long tail stretches to the right, toward large values. Negative means the long tail stretches to the left. A value near zero indicates a roughly symmetric shape, like a normal bell.
solid answer
~50 sSample skewness is the third standardised moment: you convert each observation to a z-score by subtracting the sample mean and dividing by the sample standard deviation, cube those z-scores, and average them. Cubing preserves sign, so observations far above the centre contribute large positive terms and observations far below contribute large negative ones. Whichever side has the longer, thinner tail dominates the sum. Positive skewness therefore means a long right tail — time-on-page, where most sessions are short but a few run for an hour, comes out clearly positive. Negative skewness means a long left tail, such as exam scores bunched near the maximum with a few very low ones. A symmetric shape gives a value near zero. Because it is built from z-scores, skewness is dimensionless: rescaling the data from seconds to minutes leaves it unchanged.
go deeper
Be ready to state the direction rule without hesitating: positive means a long right tail, negative a long left tail, near zero roughly symmetric. Naming a familiar right-skewed measurement such as session duration seals it.
Explain the mechanics: z-scores cubed and averaged, why cubing preserves sign, and why dividing by the standard deviation makes the result unit-free and comparable across variables.
Show judgment about reliability. Talk about how one extreme record moves a cubed sum, how sample size changes how much you trust the number, and when you would report shape at all rather than a bare summary.
Own the framing question: what does asymmetry cost the decision downstream, and when is a shape statistic worth putting into a standard reporting template versus leaving to ad-hoc analysis.
## The definition Skewness is the **third standardised moment** — a single dimensionless number describing how asymmetric a distribution is around its own centre. For a sample `x1 ... xn` with mean `xbar` and standard deviation `s`, the simplest (population-style, or *biased*) sample skewness is `g1 = (1/n) * sum( ((xi - xbar) / s)^3 )` Read that inside out. `(xi - xbar)` is a deviation from the centre. Dividing by `s` turns it into a z-score, so the quantity is expressed in standard deviations rather than in dollars or seconds. Cubing does two things at once: it **preserves the sign** (a negative deviation cubed stays negative) and it **amplifies distance** (a point 3 SD out contributes 27 units; a point 1 SD out contributes 1). Averaging those cubed z-scores gives one number whose sign reports which side of the centre carries the far-out points. Most reporting conventions use a small-sample-corrected version, the adjusted Fisher-Pearson coefficient `G1 = sqrt(n(n-1)) / (n-2) * g1` which inflates `g1` slightly on small samples. The two agree closely once `n` is large, and both carry the same sign, so the interpretation below is identical. ## Reading the sign - **Positive skewness (right-skewed, positively skewed):** the tail on the high side is long and thin. Session durations, house prices, income, and file sizes all behave this way — a floor at zero, a dense cluster of ordinary values, and a handful of very large ones that have nowhere to hide. - **Negative skewness (left-skewed):** the tail on the low side is long and thin. Think of a test where nearly everyone scores between 85 and 100 and a few score 30. - **Near zero:** consistent with symmetry. A normal distribution has population skewness exactly 0, and so does any symmetric distribution whose third moment exists. A useful mnemonic: the sign names the direction the *tail* points, not where the bulk of the data sits. Right-skewed data has *most* of its observations on the left, packed together — the name comes from the sparse tail, not the dense body. ## Zero skewness does not prove symmetry This is the classic follow-up. Symmetry implies zero skewness, but the converse fails: you can build asymmetric distributions whose positive and negative cubed deviations happen to cancel exactly. Zero is *consistent with* symmetry, not proof of it. Treat skewness as one summary of shape, alongside a look at the actual distribution, rather than a verdict. ## Units, scaling, and reflection Because every deviation is divided by `s`, skewness is **unit-free**. Multiply every value by a positive constant, or add a constant to every value, and skewness does not move: converting session durations from seconds to minutes, or from a timestamp offset to elapsed time, leaves the number identical. This is exactly why it is standardised — it makes shape comparable across variables measured on completely different scales. Multiplying by a *negative* constant reflects the distribution: the magnitude stays the same and the sign flips. A right-skewed variable becomes left-skewed when you negate it. ## Sensitivity The cube gives distant points enormous leverage. A single observation 6 SD from the mean contributes 216 units to the sum while a typical point contributes about 1. So sample skewness can swing substantially when one extreme record enters or leaves the data, and it is a noisy statistic on small samples. On tens of thousands of rows the sign is usually stable and worth reporting; on twenty rows, treat both the size and the sign with suspicion. ## Rough magnitude guidance There is no universal threshold, and interviewers do not expect one, but practitioners commonly treat |skewness| below about 0.5 as near-symmetric, 0.5 to 1 as moderately skewed, and above 1 as strongly skewed. Some reference values help calibrate: an exponential distribution has population skewness exactly 2; any symmetric distribution has 0. Quote these as orientation, not as a decision rule. ## What to do with the answer Knowing the sign and rough size tells you that a symmetric-shape assumption is questionable and that a plain mean-plus-SD summary is hiding structure. It is also the trigger for considering a transform on strictly positive data — taking logs pulls in a long right tail and can bring skewness close to zero, which is why heavily right-skewed measurements are so often summarised or modelled on a log scale.
- Does a sample skewness of exactly zero prove the distribution is symmetric?No. Symmetry forces zero skewness, but not the reverse. An asymmetric distribution can have positive and negative cubed deviations that cancel exactly, giving zero. Zero is consistent with symmetry and nothing stronger, so treat it as one summary of shape rather than a verdict.
- Why is skewness divided by the cube of the standard deviation rather than reported raw?The raw third moment has the units of the data cubed, so it changes if you switch from seconds to minutes and cannot be compared across variables. Dividing by `s^3` cancels the units and produces a dimensionless number that is unchanged by shifting or positively rescaling the data.
- What happens to skewness if you multiply every observation by -1?The magnitude is unchanged and the sign flips. Negating reflects the distribution about zero, so a long right tail becomes a long left tail. A right-skewed variable with skewness +1.4 becomes left-skewed with skewness -1.4.
- How stable is sample skewness when one very large value is added to the data?Not very. Deviations are cubed, so a point 6 SD out contributes roughly 216 times what a 1 SD point contributes. On a small or moderate sample, one extreme record can visibly move the estimate, which is why the statistic should be read with its sample size in mind.
Picture a crowd photo: skewness reports which side of the room the stragglers are standing on, not where the crowd itself is.
saying these in an interview costs you the question
- Says positive skew means most values are large
- Reads skewness as where the peak sits, ignoring the tail
- Treats a skewness of zero as proof of symmetry
- Claims skewness carries the units of the data
- Thinks converting seconds to minutes changes the skewness