What information does the sign test use from each matched pair?
answer
- which one won, nothing more
- ties contribute nothing and are dropped
- coin flips under the null
- exact binomial with p equal to one half
- weakest assumptions, weakest power
basics
~10 sOnly the direction: which member of the pair was larger. Magnitudes are discarded, ties are dropped, and the count of positive pairs is compared with a binomial distribution with success probability one half.
solid answer
~50 sFor each pair you record a plus if the first member was larger and a minus if the second was, discarding ties. If there are n non-tied pairs, the number of pluses under the null follows a `Binomial(n, 0.5)` distribution, so the p-value comes straight from binomial tail probabilities — no approximation needed. The null being tested is that a randomly chosen pair is equally likely to go either way, which is exactly the statement that the median of the paired differences is zero. Because it assumes nothing about symmetry or shape, it is the fallback when the differences are wildly asymmetric or when a pair only supports a judgement of which is better, with no measurable gap. The price is power: its asymptotic relative efficiency against a mean-based paired comparison is `2/pi`, about 0.64 under normality, and two-thirds relative to a signed-rank comparison.
go deeper
Know that it counts only which member of each pair was larger and compares that count with coin flips. Recognise it as the simplest paired test, used when magnitudes are unavailable.
Explain the mechanics: drop ties, count the pluses, refer to a Binomial(n, 0.5) null, and read the exact tail probability. State that the hypothesis is a zero median difference with no symmetry assumption.
Show the judgment about when discarding magnitude is right — untrustworthy measurement, grossly asymmetric differences, direction-only judgements — and check the minimum sample the binomial discreteness allows before committing to the design.
Own the tradeoff between assumption safety and power across a programme of studies, and set expectations that a method needing fewer assumptions also needs a larger sample budget to be worth running.
## The simplest rank-based test The sign test is what remains after you strip a paired comparison down to its bones. For each matched pair — the same subject before and after, or two treatments applied to the same unit — you ask one question: which one was bigger? Record a plus or a minus. Discard the pairs where they were equal, reducing n. If `S` is the number of pluses among n non-tied pairs, then under the null hypothesis ``` S ~ Binomial(n, 0.5) ``` and the two-sided p-value is the binomial probability of a count at least as far from `n/2` as the one observed. That exactness — no asymptotic approximation, no ranking, no reference table beyond the binomial — is part of the appeal. ## What the null actually says The null hypothesis is that for a randomly selected pair, `P(first is larger) = P(second is larger)`. Equivalently, the **median of the paired differences is zero**. Notice what is *not* assumed: - No normality. - No symmetry of the differences. - No requirement that the magnitudes be measurable, comparable, or even meaningful. It needs only that pairs be independent of one another and that the direction within each pair be recorded consistently. This is about the weakest assumption set of any test in common use, and it is the reason the sign test survives situations where everything else is arguable. ## When it earns its place 1. **Only direction is observable.** A judge says recipe A is better than recipe B but cannot put a number on how much better. There is no difference to rank, so a method that ranks magnitudes has nothing to work with. 2. **The differences are grossly asymmetric.** A test that assumes the differences are symmetric about their centre can reject because of the asymmetry rather than the centre. The sign test makes no such assumption, so its rejection means what it says. 3. **The measurement scale is untrustworthy but the ordering is not.** Instrument drift, subjective scoring and censored readings can distort magnitudes while leaving the within-pair comparison sound. ## The cost Throwing away magnitude is expensive. Asymptotic relative efficiency against a mean-based paired comparison under normality is `2/pi`, about 0.64: you need roughly 1.57 times as many pairs for the same power. Against a signed-rank comparison, which uses the ordering of magnitudes as well as the signs, the efficiency is two-thirds. That is a large penalty by the standards of this family, where the usual cost of going rank-based is a few percent. The discreteness of the binomial also bites at small n. With six non-tied pairs, the most extreme possible outcome — all six in one direction — gives a two-sided p-value of `2 * (1/2)^6 = 0.03125`, which is the smallest p-value obtainable at all. With five pairs the floor is `0.0625`, above a conventional 0.05 threshold, so no data whatsoever can reach significance. Knowing that a minimum sample size exists before a design can possibly reject is a genuinely useful piece of practical statistics. ## Handling ties Pairs with no difference carry no directional information and are conventionally dropped, with n reduced accordingly. This is exact but wasteful: a study where most pairs tie has very little effective sample, and reporting the number of ties is essential context, because a hundred pairs of which eighty tied is a much weaker study than twenty pairs with no ties. ## Relationship to the other paired methods Think of the paired tests as a ladder of how much of the data they use: - **Sign test** — directions only. - **Signed-rank test** — directions plus the ordering of the magnitudes. - **Mean-based paired comparison** — directions plus the actual magnitudes. Each step up uses more information and buys power, and each step up buys it with an additional assumption about how much you trust the magnitudes. The right rung is the highest one whose assumptions your data actually support. ## Reporting Give the number of pluses, the number of minuses, the number of ties, and the exact binomial p-value. Add the median difference where the magnitudes exist and are meaningful, so the reader knows the direction and rough size and not only that a difference was detected.
- When would you prefer the sign test over a signed-rank comparison of the same pairs?When the magnitudes are untrustworthy or unavailable — a judge can say A beat B but not by how much — or when the paired differences are so asymmetric that a symmetry assumption is indefensible. In exchange you accept roughly a one-third efficiency loss, so it is a deliberate trade of power for assumptions.
- What null hypothesis does the sign test actually test?That a randomly chosen pair is equally likely to go either way, which is exactly the statement that the median of the paired differences is zero. It says nothing about means and needs no assumption of symmetry or of any particular distributional shape, only independence between pairs.
- Why can a sign test on five non-tied pairs never reach a 0.05 threshold?The count of pluses is binomial with success probability one half, so the most extreme outcome — all five in one direction — has two-sided probability 2 times (1/2)^5, which is 0.0625. No data set of that size can produce anything smaller, so the design is underpowered by construction before a single observation is collected.
It is a show of hands. You learn how many people preferred each option, never how strongly any of them felt.
saying these in an interview costs you the question
- Thinks the sign test uses the sizes of the differences
- Counts tied pairs as evidence for the null
- Claims it has the same power as a signed-rank comparison
- Says it tests the mean difference rather than the median
- Ignores that tiny samples cannot reach significance at all