When would you control false-discovery rate with Benjamini-Hochberg instead of family-wise error?
answer
- any error versus what share of errors
- screening list versus single verdict
- sorted p-values against a rising line
- step-up: last crossing, reject below
- the (i/m)q staircase at q = 0.10
basics
~20 sUse false-discovery-rate control when you are screening thousands of hypotheses and want a bounded proportion of false hits among your discoveries rather than near-certainty of none. Family-wise control is right when a single false claim is costly.
solid answer
~50 sThe two criteria make different promises. Family-wise error control bounds the probability of *any* false rejection; false-discovery rate control bounds the expected *proportion* of false rejections among the hypotheses you reject. With 20,000 gene-expression tests, Bonferroni's threshold is `0.05/20000 = 0.0000025` and essentially nothing survives, so a screening study returns an empty list. Benjamini-Hochberg at `q = 0.10` sorts the p-values ascending, finds the largest index k with `p(k) <= (k/m) q`, and rejects everything up to k. If it returns 200 genes, roughly 10 percent of those are expected to be false leads — acceptable, because each hit goes to a cheap follow-up assay that will expose the duds. Choose false-discovery control when discoveries are screened downstream and misses are expensive; choose family-wise control when a single false positive becomes a published claim, a regulatory approval, or a shipped decision.
go deeper
Recall the two names and what each bounds: family-wise error is the chance of any false positive, false-discovery rate is the share of your positives that are false.
Be ready to run the procedure: sort the p-values, compare each to i over m times q, take the largest passing index, and reject everything at or below it.
Demonstrate the choice as a cost argument - what a single false claim costs versus what a missed effect costs, and whether hits get validated downstream - not as a preference between formulas.
Own the contract with the reader: decide which criterion the organisation reports under for which class of decision, and make sure the level, procedure and family size travel with every published result.
## Two different promises Start by writing down what each criterion controls, using V for the number of false rejections and R for the total number of rejections. - **Family-wise error rate (FWER)** = `P(V >= 1)`. The probability that you make even one false claim. - **False-discovery rate (FDR)** = `E[V / R]`, with the ratio defined as 0 when `R = 0`. The expected fraction of your announced discoveries that are wrong. These are not stricter and looser versions of the same statement; they answer different questions. FWER control at 0.05 says: with 95 percent probability this entire report contains no error at all. FDR control at q = 0.10 says: on average, about one in ten of the things on this list is a false lead. If you make no discoveries, FWER is satisfied trivially and FDR is satisfied trivially, which is precisely why FWER procedures degenerate into returning nothing when m is huge. Note also that FDR control is weaker in the following exact sense: when *every* null is true, R consists entirely of false positives, so V/R is 1 whenever anything is rejected and FDR equals FWER. Under that global null, controlling FDR at q also controls FWER at q. As soon as some effects are real, the two diverge and FDR is the more permissive criterion. ## The Benjamini-Hochberg procedure Given m p-values and a target rate q: 1. Sort ascending: `p(1) <= p(2) <= ... <= p(m)`. 2. Compute the staircase of thresholds `(i/m) * q` for `i = 1..m`. 3. Find the **largest** k such that `p(k) <= (k/m) * q`. 4. Reject `p(1)` through `p(k)` — all of them, including any that individually sit above their own threshold. Step 4 is what makes this a **step-up** procedure: you scan from the bottom of the list for the last crossing, then reject everything above it. Compare Bonferroni, which is the flat line `q/m` for every i, and Holm, whose thresholds depend on position but which stops at the *first* failure. Benjamini-Hochberg's line rises from `q/m` at the smallest p-value all the way to `q` at the largest — enormously more permissive at the far end. With m = 20,000 and q = 0.10, the smallest p-value is compared to `0.10/20000 = 0.000005` and the thousandth-smallest to `(1000/20000)*0.10 = 0.005`. A gene with p = 0.004 is unreachable for any family-wise procedure but comfortably inside the staircase if enough genes ahead of it are also small — the procedure adapts to how much signal the data appear to contain. ## Assumptions Benjamini-Hochberg controls FDR at q when the tests are independent, and also under a positive-dependence condition (positive regression dependence on the subset of true nulls), which covers many practical cases like correlated expression measurements or correlated features. Under arbitrary dependence, the Benjamini-Yekutieli variant restores control by dividing q by `1 + 1/2 + 1/3 + ... + 1/m`, which for m = 20,000 is roughly a factor of 10 — a heavy price, and one reason people check whether the positive-dependence story is plausible before reaching for it. ## Choosing between them Ask what a single false positive costs, and what happens to a discovery after you announce it. Use **family-wise** control when: - One confirmatory claim is being made and it will be acted on directly — a regulatory endpoint, a launch decision, a published causal claim. - The number of tests is small enough that the power cost is bearable. - You need a guarantee that survives arbitrary dependence with no argument about correlation structure. Use **false-discovery-rate** control when: - You are screening: hundreds or thousands of candidates, and the output is a shortlist rather than a verdict. - Each hit will be validated downstream by something cheaper than the cost of missing a real effect — a follow-up assay, a confirmatory experiment, a manual review. - Returning an empty list is itself a failure, which is exactly what family-wise procedures produce at m in the thousands. The reframing that matters in an interview: false-discovery control is not a laxer standard you adopt because the strict one was inconvenient. It is a different, honestly-stated contract with the reader — *about a tenth of this list is noise* — which happens to be the appropriate contract for screening work. Announcing a 200-gene shortlist under FDR control and announcing a single approved drug endpoint under family-wise control are both rigorous; swapping the two criteria would make both wrong. ## Reporting Whatever you choose, say it explicitly: the criterion, the level, the size of the family, and the procedure. A list of adjusted p-values with no statement of what was controlled at what level is uninterpretable, because the same number means something completely different under the two contracts.
- Under Benjamini-Hochberg, why do you reject every p-value below the last crossing rather than only those under their own threshold?Because it is a step-up procedure: the largest k satisfying p(k) <= (k/m)q defines a single cutoff, and every p-value at or below p(k) is rejected by definition of that cutoff. A p-value ranked below k that sits above its own staircase point is still smaller than p(k), so it is inside the rejection region. Testing each index in isolation would not control the false-discovery rate.
- What does Benjamini-Hochberg control when every null hypothesis is true?In that case any rejection is a false one, so V/R equals 1 whenever anything is rejected and the false-discovery rate collapses to the probability of at least one rejection. Controlling it at q therefore also controls family-wise error at q under the global null. The two criteria only diverge once some effects are real.
- Your 20,000-test screen yields nothing under Bonferroni but 200 hits under Benjamini-Hochberg at q = 0.10. How do you report that?Report the shortlist with the criterion attached: 200 candidates at a 10 percent false-discovery rate, meaning roughly 20 are expected to be spurious, and state the family size and procedure. Do not present them as 200 confirmed findings. The list is an input to validation, and the follow-up assay is what turns a candidate into a claim.
Family-wise control is a smoke alarm that must never sound falsely; false-discovery control is a security screening line that accepts a known share of false alarms because each one is cheap to check.
saying these in an interview costs you the question
- Calls false-discovery control just a looser alpha
- Says Benjamini-Hochberg bounds the chance of any error
- Applies the staircase without sorting the p-values
- Rejects only p-values under their own threshold
- Treats a false-discovery shortlist as confirmed findings
- Uses Bonferroni on twenty thousand screening tests