How does lowering the significance level from 0.05 to 0.01 affect Type II errors at a fixed sample size?
answer
- one threshold, two overlapping distributions
- the bar moves outward
- 1.96 becomes 2.576 in a two-sided z-test
- misses rise as false alarms fall
- only more evidence lowers both
basics
~20 sLowering the significance level raises the Type II error rate. Demanding stronger evidence shrinks the rejection region, so with the same data and the same true effect, more real effects fall short of the threshold and go undetected.
solid answer
~50 sLowering the significance level makes rejection harder, so the Type II error rate goes up. Concretely, for a two-sided z-test the critical value moves from about `1.96` to about `2.576`: the rejection region shrinks, and results that would have cleared the old bar now land in the non-rejection region. With the sample size, the variability and the true effect all held fixed, every one of those results is a real effect the test now misses, so beta rises. The trade-off is structural, not a flaw in any particular test — at fixed evidence you can move the threshold in one direction or the other, but you cannot lower both error rates by moving it. The only things that lower alpha and beta together are changes to the evidence itself: more observations, less measurement noise, a tighter design, or a genuinely larger effect.
go deeper
Remember the direction of the trade: tighter significance level means fewer false alarms and more missed effects. Do not claim a stricter threshold improves everything.
Draw the two sampling distributions and the single threshold between them, and explain why one line cannot avoid cutting into both. Name the critical values that go with 0.05 and 0.01 in a two-sided z-test.
Show that you spot the real-world consequence: a team that tightens the threshold after a noisy quarter has bought silence, not rigour. Redirect the conversation to sample size, variance reduction and design.
Be prepared to argue where the threshold belongs for a class of decisions and to defend the misses it implies. The interesting question is who bears the cost of the errors you deliberately chose to accept.
## What the significance level actually controls A hypothesis test converts data into a test statistic and rejects the null hypothesis when that statistic falls into a **rejection region**. The significance level `alpha` is what defines the boundary of that region: it is chosen so that, if the null hypothesis were true, the statistic would land in the rejection region only a fraction alpha of the time. So alpha is a knob on the threshold, and moving it moves the boundary. For a two-sided test based on a statistic that is approximately standard normal under the null, `alpha = 0.05` puts the boundary at roughly `|z| = 1.96`, and `alpha = 0.01` pushes it out to roughly `|z| = 2.576`. Nothing about the data changed; only the bar the data must clear. ## Two distributions, one boundary The clearest mental picture has two sampling distributions of the same statistic drawn on the same axis. One is the distribution assuming the null hypothesis is true, centred at zero. The other is the distribution assuming a particular real effect exists, shifted away from zero by an amount that grows with the size of the effect and shrinks with the noise and with the square root of the sample size. The threshold is a vertical line. The area of the null-centred distribution beyond the line is alpha, the false-alarm rate. The area of the shifted distribution that falls *short* of the line is beta, the miss rate. There is one line and two areas, and they move in opposite directions: - Push the line outward (smaller alpha) and the null distribution's tail beyond it shrinks — fewer false alarms — while more of the shifted distribution now sits on the non-rejection side, so beta rises. - Pull the line inward (larger alpha) and you accept more false alarms in exchange for fewer misses. At a fixed sample size, fixed variance and fixed true effect, that is the entire trade-off. The two distributions overlap, and no placement of a single line can avoid cutting into both of them. ## What actually moves both rates down The threshold is not the only thing on the diagram. The *separation* between the two distributions can change, and that is the real lever: - **More observations.** The spread of a sample mean shrinks with the square root of the sample size: `SE = s / sqrt(n)`. Narrower distributions overlap less, so at any given alpha the miss rate falls. - **Less measurement noise.** Reducing the variability in the underlying measurements, through better instruments, cleaner data or blocking out a known nuisance source, has the same narrowing effect as adding data. - **A more efficient design.** Paired or within-subject comparisons remove between-unit variability from the comparison and can shift the alternative distribution further out for the same number of units. - **A larger true effect.** Not usually under your control, but it is why a design that reliably catches a large effect can be helpless against a small one. This is why the honest answer to "we got a null result, should we just use a stricter threshold next time?" is no: strictness is not rigour, it is a re-allocation of risk from one error type to the other. ## A worked intuition Suppose a design gives a statistic that, if the real effect is what you expect, sits on average at `z = 2.2`. At `alpha = 0.05` the bar is 1.96, and a typical outcome clears it. Move the bar to 2.576 and that same typical outcome now falls short: the effect is real, present in the data at its expected magnitude, and the test reports nothing. Nothing was measured worse; the decision rule simply stopped counting that much evidence as enough. ## Where the choice comes from Because moving alpha only re-allocates risk, the choice has to come from outside the mathematics — from what each mistake costs. A setting where a false alarm triggers an expensive, hard-to-reverse commitment argues for a strict threshold; a setting where a miss is the expensive outcome and a false alarm merely triggers a cheap follow-up argues for a loose one. What is never defensible is choosing the threshold after seeing which side of it your result landed on, because the stated false-alarm rate then no longer describes the procedure you actually ran. ## Common confusions to avoid Candidates sometimes claim a stricter alpha makes a significant result "more likely to be true" in a way that costs nothing. It does buy something real — surviving results were held to a higher evidentiary bar — but the cost is paid in silence, in the real effects that never get reported. Others assume beta has a fixed value that alpha cannot touch; in fact beta depends jointly on alpha, the sample size, the variability and the size of the true effect, and quoting it without those four is meaningless.
- If moving the threshold only re-allocates risk, what actually lowers both error rates at once?Changes to the evidence rather than the rule. More observations shrink the standard error as `s / sqrt(n)`, reducing the overlap between the null and alternative sampling distributions. So does cutting measurement noise or using a paired design that removes between-unit variability. With the distributions further apart, any fixed threshold produces both fewer false alarms and fewer misses.
- Does a stricter significance level make the results that do reach significance more trustworthy?They cleared a higher evidentiary bar, so yes in that narrow sense. But the gain is not free: the same strictness sends more real effects into the non-significant pile, and those misses are invisible in the published record. Judging a stricter threshold only by what survives it ignores half of its consequences.
- Why is it illegitimate to choose the significance level after seeing the test statistic?The stated false-alarm rate describes a procedure fixed in advance. If the threshold is set to land on whichever side of the result you prefer, the actual procedure is 'reject whenever it suits me', whose real Type I error rate is far above the number reported. Fix alpha before the data, and record that you did.
It is the sensitivity dial on a smoke alarm with the hardware unchanged. Turn it down and you stop being woken by toast; you have also made it likelier that a slow smoulder goes unannounced.
saying these in an interview costs you the question
- Says a stricter threshold lowers both error rates
- Treats the Type II error rate as unaffected by alpha
- Calls a smaller significance level simply more rigorous
- Believes only sample size influences the miss rate
- Picks the threshold after seeing the test statistic