An early-stage drug safety screen is run at alpha = 0.10 rather than 0.05. How would you defend that choice?
answer
- name the null before pricing anything
- which mistake is unrecoverable
- who absorbs each cost
- a miss here advances harm
- pre-specify, and keep the screen upstream
basics
~20 sThe two errors have very different costs here. The null is that the compound is safe, so a miss advances a possibly harmful compound while a false alarm only triggers extra testing; the looser threshold buys a lower miss rate.
solid answer
~50 sStart by naming the hypotheses, because the defence depends on them: the null is that the compound shows no adverse effect, so rejecting means flagging a safety signal. Now price the two errors. A Type I error flags a compound that is actually fine — at an early screen that costs some follow-up work, and the mistake is caught downstream. A Type II error lets a genuinely harmful compound pass the screen and advance towards people, which is expensive and hard to reverse. When the miss is the costlier and less recoverable error, you buy a lower miss rate with a looser threshold: `alpha = 0.10` at the same sample size. The defence has to include two more things — that the level was fixed before the data, and that the screen sits upstream of a stricter confirmatory stage that will not inherit the loose threshold.
go deeper
Know that the significance level is a choice, not a law of nature, and that 0.05 is only a convention. You will not be asked to defend a departure yet, but do not call one wrong on sight.
Explain mechanically what raising the level to 0.10 does: the rejection region widens, false alarms become more common, missed effects less so, with the sample size unchanged.
Demonstrate the full cost argument — name the null, price both errors including recoverability and who pays, then state the guardrails of pre-specification and a stricter confirmatory stage.
Own the position that error thresholds encode institutional risk appetite, and be ready to defend that appetite to regulators, clinicians or executives who do not read statistics.
## The question behind the question Interviewers ask this to find out whether a candidate treats `0.05` as a scientific constant or as a decision parameter. It is the latter. The 5% convention is a historical default with no derivation behind it, and defending a departure from it is ordinary applied statistics, not a violation. ## Step one: name the null Nothing about error costs can be reasoned about until the hypotheses are stated, because the labels attach to them. In a safety screen the null hypothesis is typically "this compound produces no adverse effect", and rejecting the null means raising a safety flag. That fixes the two mistakes: - **Type I error** — flagging a compound that is in fact safe. Consequence: additional assays, a delay, some wasted budget. - **Type II error** — clearing a compound that in fact causes harm. Consequence: a harmful candidate advances through the pipeline, potentially towards human exposure. Swap the null and every conclusion inverts, which is why a candidate who launches straight into cost talk without saying which hypothesis is the null has not actually answered. ## Step two: price the errors, including recoverability A cost comparison has three components, and candidates usually give only the first. 1. **Magnitude.** What does each mistake cost in money, time, health or reputation? 2. **Recoverability.** Is there a later stage that catches the mistake? A false alarm at a screen is usually corrected by the very next experiment. A miss at a screen is not corrected at all, because nothing downstream re-examines what was already discarded — the compound simply moves on unflagged. 3. **Who bears it.** A false alarm costs the sponsor. A missed harm is borne by trial participants and, later, patients. That asymmetry is the strongest part of the argument and belongs in the answer. On all three, the miss dominates in an early safety screen. That is the case for `alpha = 0.10`: at a fixed sample size, a looser threshold shrinks the miss rate, and the extra false alarms are absorbed by a cheap, self-correcting follow-up step. ## Step three: say what the looser threshold does not buy A strong answer volunteers the limits. - **A looser alpha does not compensate for a small sample.** Loosening the threshold trades one error for the other; it does not add information. If the design is too small to see the harms that matter, `alpha = 0.10` is a fig leaf and the honest fix is a bigger or better-designed study. - **The loose threshold must not propagate.** A screen exists to feed a stricter confirmatory stage. If the flagged compounds are treated as established findings, the extra false alarms you deliberately bought become false conclusions. - **The level must be pre-specified.** Choosing 0.10 after seeing a result at `p = 0.08` is not a cost argument, it is threshold-shopping, and it invalidates the stated false-alarm rate. Write the level and the reasoning down before the data exist. ## The mirror case To show the reasoning generalises, describe when you would move the other way. Push towards `alpha = 0.01` when a false alarm triggers the expensive, irreversible action and a miss is cheap or recoverable: a decision to halt a production line, recall a product, or commit a large capital spend on the strength of one analysis. There the institution absorbs an enormous cost for acting on a signal that is not real, and delay merely means waiting for more data. Same framework, opposite conclusion — which is precisely the point. ## How to structure the spoken answer A clean thirty seconds: state the null; state what each error means in this specific setting; compare magnitude, recoverability and who pays; conclude that the miss dominates, so the threshold loosens; then add the two guardrails, pre-specification and a stricter confirmatory stage. That structure works for any "why this alpha?" question in any domain, which is why interviewers like it. ## What weak answers sound like "0.05 is the standard, so 0.10 is sloppy" treats a convention as a law. "We used 0.10 because the sample was small" mistakes a threshold for evidence. "We wanted the result to be significant" is the failure mode the pre-specification rule exists to prevent, and saying it out loud usually ends the discussion.
- In what situation would you argue the other way, for a threshold of 0.01 instead?When a false alarm triggers the expensive and irreversible action and a miss is cheap or recoverable — halting a production line, issuing a recall, or committing major capital on one analysis. There, acting on a signal that is not real costs the organisation far more than waiting for more evidence, so the threshold should be strict.
- Can a looser significance level substitute for a sample size that is too small?No. Loosening the threshold moves risk from misses to false alarms; it adds no information. If the design cannot detect the harm magnitudes that matter, a looser alpha just relabels an undersized screen as a permissive one. The real remedies are more units, less measurement noise, or a design that removes nuisance variability.
- What has to be true downstream for a deliberately loose screening threshold to be safe?The screen must feed a stricter confirmatory stage that re-tests the flagged compounds independently, and the flags must be reported as candidates rather than findings. If a downstream team treats the screen's output as established, every extra false alarm you accepted turns into a false conclusion that nothing corrects.
saying these in an interview costs you the question
- Treats 0.05 as a fixed scientific standard
- Argues about costs without naming the null hypothesis
- Uses a looser threshold to compensate for a small sample
- Chooses the level after seeing the result
- Assumes false positives are always the costlier error