How does the Wilcoxon signed-rank test use the differences within matched pairs?
answer
- one difference per pair, then order them
- rank the sizes, keep the signs aside
- zeros are usually dropped first
- add up the ranks of the positive differences
- symmetry is the assumption that buys the median
basics
~20 sIt takes each pair's difference, drops zeros, ranks the absolute differences from smallest to largest, and sums the ranks belonging to positive differences. Large or small sums indicate the differences are not centred at zero.
solid answer
~50 sFor each matched pair you compute a difference `d = first - second`. Pairs with `d = 0` are usually dropped and the sample size reduced. You then rank the remaining `|d|` values from 1 upward, using midranks for ties, re-attach each difference's sign, and sum the ranks carrying a positive sign to get `W+`. Under the null hypothesis the differences are symmetrically distributed about zero, so each sign is equally likely and `W+` has mean `n(n+1)/4`; a value far from that in either direction is evidence against the null. In a paired taste test where every judge scores both recipes, this uses two things the sign test throws away: which recipe won each pairing *and* how the sizes of the gaps line up. The price is the symmetry assumption — the test targets a symmetric centre of the differences, not any arbitrary median.
go deeper
Know that it is the rank method for paired or before-and-after measurements: one difference per pair, ranked by size, with the signs kept. Recognise it as the paired counterpart of a rank comparison of two independent groups.
Walk through the mechanics: differences, drop zeros, rank absolute values with midranks, sum the positive ranks, compare with mean n(n+1)/4. Say that the null is symmetry of the differences about zero.
Demonstrate that you check symmetry before claiming a median difference, decide deliberately how to handle zeros, and report the pseudomedian as the effect estimate rather than a bare p-value.
Own the design question: whether pairing is genuinely available and worth the logistics, and what the team commits to as the reported estimand for paired outcomes so results stay comparable across studies.
## The procedure, step by step Suppose ten judges each taste two recipes and score both, so every judge supplies a matched pair. 1. Form the within-pair difference `d_i = score_A - score_B` for each judge. Pairing removes the judge-to-judge variation, which is exactly why paired designs are powerful. 2. Discard pairs with `d_i = 0` and reduce n accordingly. (Pratt's variant instead ranks the zeros and then drops their contribution; it is less common but less wasteful.) 3. Rank the absolute differences `|d_i|` from 1 for the smallest to n for the largest, assigning midranks to ties. 4. Re-attach the sign of each difference to its rank. 5. Sum the ranks with a positive sign to get `W+` (and the negatives to get `W-`). Note `W+ + W- = n(n+1)/2`. ## The null distribution Under the null hypothesis, the distribution of the differences is symmetric about zero. Symmetry means that for each pair, the sign attached to a given magnitude is a coin flip independent of that magnitude. So the null distribution of `W+` is that of a sum of `n(n+1)/2` split randomly by independent signs: ``` mean(W+) = n(n + 1)/4 var(W+) = n(n + 1)(2n + 1)/24 ``` For small n the exact distribution is enumerated over all 2^n sign patterns; for larger n the normal approximation with those moments is used, with a correction when ties are present. Either tail is evidence: `W+` far above its mean says positive differences dominate, far below says negative ones do. ## What the test is really testing The careful statement is that the differences are symmetrically distributed about zero. Under symmetry the mean, the median and the centre of symmetry all coincide, which is why the test is usually described as testing whether the median difference is zero. The estimand that goes with it is the **pseudomedian**: the median of all pairwise averages `(d_i + d_j)/2` over pairs of differences, i and j including i = j. This is the Hodges-Lehmann estimator for the one-sample problem and is the number you should report as the point estimate alongside the p-value. If the differences are strongly skewed, symmetry fails, and a rejection can be driven by asymmetry rather than by a shift away from zero. That is the single most useful caveat to volunteer in an interview. ## Why it is stronger than counting winners The test uses the *ordering of magnitudes*, not just the directions. Imagine seven judges narrowly preferring recipe A and three strongly preferring recipe B. Counting winners alone points at A. The signed-rank test notices that the three B preferences carry the three largest absolute differences and therefore the largest ranks, so `W+` is pulled back toward its null mean. Whether that is what you want depends on whether the magnitudes are trustworthy — see the assumption above. At the same time it does not use the raw magnitudes: a difference of 0.4 and a difference of 400 contribute the same rank if they hold the same position in the ordering. That is the source of the outlier robustness. ## Assumptions 1. **Pairs are independent of one another.** Within a pair the two measurements are deliberately dependent — that is the design — but judge 1's pair must not influence judge 2's. 2. **The differences are measured on at least an interval-like scale**, enough that comparing the sizes of two differences is meaningful. For a purely ordinal within-pair judgement of the form A is better than B, with no magnitude, the signed-rank ordering of magnitudes is not defensible. 3. **Symmetry of the differences under the null.** This is what buys the clean median interpretation. ## Common mistakes - Applying it to two independent groups. Signed-rank is for paired data; independent samples call for a rank-sum comparison instead. - Treating the drop-the-zeros step as harmless. Many zeros mean a real loss of information and of power, and n shrinks with them. - Ranking the signed differences instead of the absolute differences. The ranking is on `|d|`; the sign is re-attached afterwards. - Claiming the test needs no assumptions. It needs independence across pairs and symmetry for the median reading. ## Reporting Give `W+` (or the smaller of `W+` and `W-`, which is what many tables use), the number of non-zero pairs, the p-value, and the Hodges-Lehmann pseudomedian as the point estimate of the typical difference. Say explicitly whether you dropped zeros and how many.
- What breaks if the paired differences are strongly skewed?The symmetry assumption fails, so a small p-value may reflect asymmetry rather than a centre away from zero, and the median interpretation stops being clean. Options are to work on a scale where the differences are more symmetric, or to fall back on a test of direction only, which assumes nothing about symmetry.
- How should zero differences be handled?The standard procedure drops them and reduces n before ranking, which is simple but discards information and power when zeros are common. Pratt's variant ranks all absolute differences including the zeros and then excludes the zero ranks from the sum, keeping their effect on the ranking. Always report how many zeros there were.
- What point estimate accompanies a Wilcoxon signed-rank p-value?The pseudomedian: the median of all pairwise averages of the differences, known as the one-sample Hodges-Lehmann estimator. It is the location parameter the test is built around and equals the median difference when the differences are symmetric. Reporting a p-value with no estimate of the typical difference is incomplete.
saying these in an interview costs you the question
- Uses the signed-rank test on two independent groups
- Ranks the signed differences instead of the absolute ones
- Never mentions the symmetry assumption
- Treats dropping zero differences as costless
- Says the test uses the raw sizes of the differences