What does a team lose by defaulting to rank-based tests whenever the data look skewed?
answer
- the question changes, not just the method
- sums and orderings can disagree
- roughly five percent more sample at the normal
- shapes must match for a location claim
- the estimand comes before the test
basics
~20 sThey quietly change the question. A rank test asks how often one group is higher, not by how much, so it can be significant while the total the business banks is unchanged — and it costs about five percent efficiency.
solid answer
~50 sThe first cost is the hypothesis itself. Ranks test stochastic ordering — whether a value from one group tends to exceed a value from the other — while most business quantities are sums or averages: revenue booked, minutes served, items shipped. Those can disagree, so a rank test can hand you a clean p-value about a quantity nobody acts on. The second cost is efficiency: on genuinely normal data the asymptotic relative efficiency against a mean-based comparison is `3/pi`, about 0.955, so you need roughly five percent more data for the same power; it never drops below 0.864 for any continuous distribution and exceeds 1 for heavy tails. Third, a rank rejection is a location statement only when the groups have similar shapes, which is exactly what skewed, heteroscedastic data violate. Name the estimand first, then pick the test.
go deeper
Know that a rank test answers how often one group is higher rather than by how much, and that using it is a change of question, not just a change of tool.
Be able to quantify the efficiency cost — about five percent more data at the normal, with a floor of 0.864 and gains on heavy tails — and to explain why equal shapes are needed for a location claim.
Show that you pre-specify the estimand and the test, investigate the source of the skew instead of routing around it, and refuse to let a histogram silently swap the hypothesis mid-analysis.
Own the standard: which estimand the organisation decides on, what the analysis plan must record before data are seen, and how disagreements between a rank result and a total are escalated rather than quietly resolved.
## The policy under scrutiny "If the data look skewed, use a rank test." It sounds cautious and is often harmless, but as a standing default it substitutes a shape diagnostic for a decision about what you are trying to estimate. Three separate costs follow. ## Cost 1: the estimand changes silently A rank comparison of two independent samples tests whether `P(X > Y) = 1/2` — how often one group's value exceeds the other's. Under an equal-shape assumption that is also a statement about a shift in location. Under no assumption it is a statement about ordering, full stop. Most operational quantities are not orderings. Revenue is a sum. Latency budgets are stated as a mean or a specific quantile. Capacity plans multiply an average by a headcount. On heavy-tailed data these can move in opposite directions to the ranks: a change that makes the bulk of users slightly worse while making a small tail dramatically better can leave the rank comparison significant in one direction and the total in the other. Whichever way the disagreement goes, the analysis has answered a question the decision does not use. The fix is to name the estimand first: "we act on total revenue, so we need an inference about the mean", or "we act on the typical user's experience, so ordering is exactly right". The test follows from the estimand, not from a histogram. ## Cost 2: efficiency, quantified Asymptotic relative efficiency (ARE) compares the sample sizes two tests need to reach the same power against the same shrinking alternative. For the rank-sum comparison against the mean-based one: - Under a true normal distribution, ARE = `3/pi` which is about 0.955. The rank test needs roughly `1/0.955`, about five percent, more observations. - Across all continuous distributions the ARE is bounded below by 0.864. The worst case is bounded — you cannot be catastrophically penalised. - For heavy-tailed distributions the ARE exceeds 1, sometimes far above it: with contaminated or long-tailed data the rank test is *more* efficient, sometimes dramatically. So the pure power argument is mild and often favours ranks. Quote these numbers rather than a vague "ranks lose power": the interviewer is checking whether you know the cost is about five percent at the normal, not fifty. ## Cost 3: interpretation and reporting Skewed data are frequently also heteroscedastic — the groups differ in spread as well as in centre. That is precisely the case where a rank rejection cannot be read as a location shift, because the null being rejected is equality of the whole distributions. A team that defaults to ranks and then reports "the median improved" is making a claim the test does not support. When shapes plainly differ, the Brunner-Munzel test targets `P(X < Y)` without an equal-shape assumption and is the more defensible choice. Related practical costs: - **Effect sizes get harder to communicate.** A probability of superiority of 0.58 is a real quantity but few stakeholders can convert it into a plan. Rank-biserial correlations fare no better in a business review. - **Very small samples give lumpy p-values.** Exact rank distributions are discrete; with n = 5 per group the smallest achievable two-sided p-value may already sit above a conventional threshold, so no result can reach significance regardless of the data. - **Heavy ties erode power.** Coarse ordinal scales generate many ties, needing tie corrections and blunting the test. ## What a strong answer proposes instead 1. **Fix the estimand before looking at the data.** Write it in the analysis plan: the mean, a specified quantile, or the probability of superiority. 2. **Investigate the skew rather than routing around it.** Is it a genuine heavy tail, a mixture of two populations, or data-entry damage? The three call for different responses, and only one of them is a test choice. 3. **Pre-specify.** Choosing the test after seeing which one gives the nicer p-value is a garden of forking paths, and a rank-versus-mean switch is one of the easiest forks to walk down unnoticed. Write the rule down in advance and state it in the report. 4. **Report both when the estimand is genuinely ambiguous** — but say up front which one the decision rests on, so a disagreement between them becomes a finding rather than a choice made after the fact. ## The one-line version A rank test is not a safer version of a mean test; it is a test of a different hypothesis with a mild power difference. Defaulting to it on the basis of a histogram trades a question you care about for a question that is easier to satisfy assumptions for.
- How much power does a rank comparison actually cost on truly normal data?The asymptotic relative efficiency against the mean-based comparison is 3/pi, about 0.955, so you need roughly five percent more observations for the same power. It never falls below 0.864 for any continuous distribution, and it rises above 1 for heavy-tailed data, where the rank test is the more efficient of the two.
- How do you decide which test to run before seeing the data?Start from the decision. Name the quantity the organisation will act on — total revenue, a latency quantile, the probability one variant is better — write it into the analysis plan along with the test that estimates it, and record how outliers and data errors will be handled. Then the shape of the histogram cannot silently change the hypothesis.
- What is the risk of choosing between a rank test and a mean-based test after inspecting the results?The real error rate no longer matches the nominal one, because you effectively took the better of two tests. That is a forking-paths problem, and the switch is easy to justify to yourself as a reasonable reaction to skew. Pre-specify the rule, or declare the second analysis explicitly as exploratory.
saying these in an interview costs you the question
- Treats a rank test as a strictly safer version of a mean-based test
- Believes rank tests lose most of their power under normality
- Chooses the test after seeing which p-value is smaller
- Reports a median improvement from a rank test on differently shaped groups
- Never connects the analysis to the quantity the business acts on