skip to content

What does the Mann-Whitney U test actually compare between two independent samples?

level: middleimportance: must knowfreq 72%

answer

  1. order across both samples, pooled
  2. count the head-to-head pairwise wins
  3. the null is one distribution, not one median
  4. medians only under equal shapes
  5. the null says P(X > Y) = 1/2

basics

~20 s

Pool both samples, rank them, and U counts how often a value from one group beats a value from the other, testing whether P(X > Y) = 1/2. It compares medians only when the two distributions share a shape.

solid answer

~50 s

The Mann-Whitney U test (identical to the Wilcoxon rank-sum test) pools the two samples, ranks all N observations together, and sums the ranks belonging to each group. `U1 = R1 - n1(n1+1)/2`, where `R1` is group 1's rank sum, and U1 counts the pairs (x, y) in which the group-1 value is larger, a tie counting as half. Under the null that both samples come from the same distribution, a randomly drawn x is equally likely to beat or lose to a randomly drawn y, so `P(X > Y) = 1/2` and U has mean `n1*n2/2`. That null is about the whole distribution: reading a rejection as a median difference requires the extra assumption that both distributions have the same shape and spread. It suits ordinal data such as 1-to-5 satisfaction ratings, where averaging stars would falsely assume the gap from 1 to 2 equals the gap from 4 to 5.

go deeper

for a junior

Know that it compares two independent groups using ranks instead of raw values, and that it suits skewed or ordinal data such as 1-to-5 ratings. Be able to say it is the same test as the Wilcoxon rank-sum test.

for a middle

Explain the machinery: pooled ranks, U as the count of pairwise wins, mean n1*n2/2 under the null, and the tie correction. State the null as P(X > Y) = 1/2 rather than as a statement about medians.

for a senior

Show that you catch the equal-shape condition before making a location claim, choose the interpretation the decision needs, and always report an effect size. Be ready to explain why an ordinal scale should not be averaged.

for a principal

Own the estimand question: whether the organisation should be deciding on stochastic ordering or on a mean the business can sum. Set the reporting standard so teams cannot quietly swap one hypothesis for the other mid-analysis.

## The statistic Take two independent samples of sizes n1 and n2. Pool all N = n1 + n2 observations, sort them, and assign ranks 1 to N, giving tied values their midrank. Let R1 be the sum of the ranks that landed in group 1. Then ``` U1 = R1 - n1(n1 + 1)/2 U2 = n1*n2 - U1 ``` U1 has a direct combinatorial meaning: it is the number of (x, y) pairs, one value from each sample, in which the group-1 value is the larger, with each tie counted as one half. There are n1*n2 such pairs in total, which is why U1 + U2 = n1*n2. Under the null hypothesis the sampling distribution of U is known exactly for small samples and is approximately normal for larger ones, with ``` mean(U) = n1*n2/2 var(U) = n1*n2*(N + 1)/12 (before any tie correction) ``` Heavy ties shrink the true variance, so implementations apply a tie correction that reduces the variance term; skipping it makes the test conservative. This test is exactly the Wilcoxon rank-sum test under a different but equivalent parameterisation — interviewers use the two names interchangeably. ## What the null hypothesis actually says The textbook null is that the two samples are drawn from the **same distribution**. Its most useful restatement is ``` P(X > Y) + 0.5 * P(X = Y) = 1/2 ``` in words: pick one observation from each group at random and either is equally likely to be the bigger one. The estimate of that quantity is `U1/(n1*n2)`, sometimes called the probability of superiority or common-language effect size. This is why U is often described as a test of **stochastic ordering** rather than of any single summary number. The alternative hypothesis in the general form is simply that the two distributions differ in a way that makes one tend to produce larger values. ## Why it is not automatically a median test This is where candidates most often slip. If the two distributions have the same shape and spread and differ only by a horizontal shift, then stochastic ordering and a median difference are the same statement, and reading a significant U as "group A tends to score higher, and its median is higher" is fair. Add a matching point estimate of the shift — the median of all pairwise differences x - y, the Hodges-Lehmann estimator — and the result is interpretable. Drop the equal-shape assumption and the equivalence breaks. Two groups can have identical medians while one is far more dispersed than the other; U can then reject, correctly, because the distributions genuinely differ, yet a claim about medians would be wrong. When shapes plainly differ, the Brunner-Munzel test targets the same `P(X < Y)` quantity without assuming a common shape and is the more honest choice. ## Where it fits naturally Ordinal outcomes are the clearest case. Satisfaction ratings on a 1-to-5 scale carry order but not distance: nothing guarantees that the step from 1 to 2 means as much as the step from 4 to 5, so an average of stars is arithmetic on labels. Ranks use only the ordering the scale really supports. Expect a great many ties on a five-point scale, which makes the tie correction and the normal approximation, rather than the exact null distribution, the relevant machinery. Strongly skewed continuous measurements are the other common case: session duration, time-to-repair, spend per user. Here you should be explicit that a rank comparison answers *how often* one group is higher, not *by how much* — and if the decision is denominated in the total or the average, that difference matters. ## Assumptions worth stating out loud 1. Observations are independent within each group and the two groups are independent of each other. Repeated measurements on the same subject break this; that design calls for a paired rank method instead. 2. The outcome is at least ordinal. 3. For a location interpretation only: similar shapes and spreads. Notably absent: normality, and any requirement that variance be equal — as long as you keep the interpretation to stochastic ordering rather than to medians. ## Reporting Quote U (or the rank sum), the sample sizes, the p-value, and an effect size. The two standard rank effect sizes are the probability of superiority `U1/(n1*n2)` and the rank-biserial correlation `2*U1/(n1*n2) - 1`, which runs from -1 to +1 and is zero when the groups are indistinguishable. A p-value with no effect size is an incomplete answer at any level above screening.

  • Two groups have the same median but very different spreads — can Mann-Whitney reject?
    Yes, and it would be right to. The null is that the distributions are identical, so a difference in shape or spread alone can produce a small p-value. The error is reporting that as a median difference. When shapes clearly differ, the Brunner-Munzel test estimates P(X < Y) without assuming a common shape.
  • How do you report an effect size for a Mann-Whitney result?
    Use the probability of superiority, U1/(n1*n2) — the estimated chance a randomly chosen value from group 1 exceeds one from group 2 — or the rank-biserial correlation 2*U1/(n1*n2) - 1, which is zero under no difference. If a shift interpretation is defensible, add the Hodges-Lehmann estimate: the median of all pairwise differences.
  • What happens to the test when a 1-to-5 rating scale produces many ties?
    Tied values take midranks, and the null variance of U must be reduced by a tie correction; without it the test is conservative and loses power. Exact small-sample null distributions assume no ties, so with a heavily tied five-point scale you rely on the tie-corrected normal approximation instead.
  • Why is the test also called the Wilcoxon rank-sum test?
    They are the same procedure. Wilcoxon framed it as the sum of ranks in one sample; Mann and Whitney framed it as the count of pairwise wins. U1 = R1 - n1(n1+1)/2 converts one to the other, so both give identical p-values. Do not confuse it with the Wilcoxon signed-rank test, which is for paired data.

It is a round-robin scoreboard: every value in one group plays every value in the other, and U is the number of matches the first group wins.

saying these in an interview costs you the question

  • Calls it a test of medians with no conditions attached
  • Says it requires no assumptions whatsoever
  • Describes it as comparing means without normality
  • Ignores ties on a short ordinal rating scale
  • Applies it to repeated measurements on the same subjects
  • Reports only a p-value with no rank effect size

context