skip to content

Errors and Power

The two ways a test can be wrong and how often it catches a real effect: alpha, beta, power and the effect size behind them. Interviewers ask because underpowered analysis is everywhere.

on this pageshow

explore

questions

14

What is the difference between a Type I and a Type II error in hypothesis testing?

level: juniorimportance: must knowfreq 88%

answer

  1. two ways a verdict disagrees with reality
  2. false alarm versus missed effect
  3. one rate is chosen, one is inherited
  4. alpha names one, beta the other
  5. convict the innocent versus free the guilty

basics

~20 s

A Type I error rejects a null hypothesis that is actually true, a false alarm whose long-run rate is alpha. A Type II error fails to reject a null hypothesis that is actually false, a missed effect whose rate is beta.

solid answer

~50 s

The two errors are the two ways a test's verdict can disagree with reality. A **Type I error** is rejecting a true null hypothesis: you declare an effect that is not there. Its long-run rate is `alpha`, the significance level you fix before collecting data, so a test run at `alpha = 0.05` produces a false alarm in about 5% of studies in which the null is genuinely true. A **Type II error** is failing to reject a false null hypothesis: a real effect is present and the test misses it. Its rate is `beta`. The asymmetry worth stating out loud is that alpha is a number you choose, while beta is a consequence you inherit from the sample size, the noise in the data and how big the true effect actually is — so beta is only defined once you name a specific alternative effect size.

go deeper

for a junior

Be ready to state both definitions cleanly and put the right label on the right cell under pressure. Say which hypothesis is being rejected in each case before reaching for the false-positive and false-negative shorthand.

for a middle

Explain the conditioning: alpha is computed assuming the null is true, beta assuming a specific alternative is true. Show why beta has no value until an effect size is named.

for a senior

Show that you translate the two errors into consequences for the decision at hand: which action each mistake triggers, who absorbs the cost, and whether it is recoverable. Definitions alone are not enough at this level.

for a principal

Own the framing that the error pair is a policy choice, not a statistical fact. Be ready to argue where an organisation should sit on that trade-off and how it defends the position to people outside the analytics team.

## The two ways a test can be wrong A hypothesis test ends in one of two verdicts: reject the null hypothesis, or fail to reject it. Reality is also in one of two states: the null is true, or it is false. Crossing those gives four cells, two of which are correct decisions and two of which are errors. | | Null is true | Null is false | |---|---|---| | **Reject the null** | Type I error (false alarm) | Correct detection | | **Fail to reject** | Correct non-detection | Type II error (miss) | **Type I error** — rejecting a null hypothesis that is in fact true. You announce an effect, a difference, an association, and there is none. Its probability, computed under the assumption that the null is true, is written `alpha` and is called the significance level. It is not discovered from the data; it is set by the analyst in advance, and the test's decision rule is built to honour it. Choosing `alpha = 0.05` means: among all the studies I run in which the null is genuinely true, I am willing for about 5 in 100 to end in a false alarm. **Type II error** — failing to reject a null hypothesis that is in fact false. The effect is real and the test does not catch it. Its probability is written `beta`. ## The conditioning is the whole point Both quantities are *conditional* probabilities, and beginners drop the condition. `alpha = P(reject the null | the null is true)`. It is emphatically not the probability that the null is true, and it is not the fraction of your published findings that are wrong — that would require knowing how often you test true nulls in the first place, which the significance level says nothing about. Similarly `beta = P(fail to reject | the null is false)`. But "the null is false" is not one state of the world; it is a whole family of them. A tiny true effect is easy to miss and a huge one is hard to miss, so beta is not a single number attached to a test. It only becomes a number once you name a specific alternative: "if the true difference is 3 points, this design misses it with probability 0.30." Any candidate who quotes a beta without naming an effect size has skipped a step. ## The courtroom framing The standard teaching analogy is a criminal trial where the null hypothesis is "the defendant is innocent" — presumed innocent, and only overturned by evidence beyond a threshold. Convicting an innocent defendant is a Type I error. Acquitting a guilty one is a Type II error. The analogy earns its keep for three reasons. First, it makes the asymmetry visible: the system does not treat the two mistakes as equally bad, which is exactly why the evidentiary bar is set high rather than at even odds. Second, it shows why "fail to reject" is not "accept" — a verdict of not guilty asserts that the evidence was insufficient, not that innocence was proven. Third, it makes obvious that you cannot drive both errors to zero: acquit everyone and you convict no innocents but free every guilty party; convict everyone and you catch every guilty party but jail the innocent. ## Why you cannot have both at zero With a fixed amount of evidence, the two error rates trade against each other. Any rule that makes rejection harder lowers the false-alarm rate and raises the miss rate, and any rule that makes rejection easier does the reverse. Setting `alpha = 0` is achievable — never reject the null, whatever the data — but it forces the miss rate to 1, which is why it is not a serious option. The only ways to push *both* rates down at once are to change the evidence rather than the threshold: collect more observations, reduce measurement noise, or use a more efficient design. ## What interviewers listen for They want the definitions stated with the conditioning intact, the labels attached to the right cell (a surprising number of candidates swap them under pressure), and an acknowledgement that alpha is a design choice while beta is a consequence. Saying "Type I is a false positive, Type II is a false negative" is a fine mnemonic, but only after you have said which hypothesis is being rejected in each case — the false-positive language is meaningless until the null is named.

  • Which of the two error rates does the analyst fix in advance, and which one is a consequence of the design?
    Alpha is fixed in advance: you declare the significance level before collecting data, and the decision rule is constructed to honour it. Beta is inherited. It falls out of the sample size, the variability in the measurements and the size of the true effect, so you can influence it through design choices but you never simply declare it.
  • Why does quoting a Type II error rate require naming a specific effect size, while a Type I rate does not?
    The null hypothesis is a single point, so the false-alarm probability can be computed once and for all under it. "The null is false" covers infinitely many worlds, from a trivial effect to an enormous one, and the miss rate differs in each. A beta is only meaningful attached to a stated alternative, such as "a 3-point difference".
  • Is 'fail to reject the null' the same as proving the null is true?
    No. Failing to reject means the evidence did not clear the threshold, which happens both when the null is true and when a real effect was too small or too noisy for this design to catch. Reporting it as proof of no effect confuses an absence of evidence with evidence of absence, and it silently ignores the Type II error rate.

Think of a criminal trial where the null hypothesis is 'presumed innocent'. Convicting an innocent defendant is a Type I error; letting a guilty one walk free is a Type II error, and no evidentiary bar removes both risks at once.

saying these in an interview costs you the question

  • Swaps the labels: calls a missed effect a Type I error
  • Says alpha is the probability that the null hypothesis is true
  • Quotes a Type II error rate without naming any effect size
  • Claims a careful enough analysis can eliminate both errors
  • Reads 'fail to reject' as 'the null is proven true'

context

open as a page

What does it mean to say a hypothesis test has 80% power?

level: juniorimportance: must knowfreq 85%

basics

~20 s

Power is the probability a test rejects the null when a specified real effect exists. 80% power means that if an effect of that size is truly present, the test detects it 80% of the time.

open as a page

How do you compute Cohen's d for a 3-point mean gap when the pooled SD is 15?

level: middleimportance: must knowfreq 74%

basics

~20 s

Divide the difference in means by the pooled standard deviation: 3 / 15 = 0.2. The two groups sit one fifth of a standard deviation apart, which Cohen's conventional benchmarks call a small effect, and their distributions overlap heavily.

open as a page

How does lowering the significance level from 0.05 to 0.01 affect Type II errors at a fixed sample size?

level: middleimportance: must knowfreq 68%

basics

~20 s

Lowering the significance level raises the Type II error rate. Demanding stronger evidence shrinks the rejection region, so with the same data and the same true effect, more real effects fall short of the threshold and go undetected.

open as a page

What four levers determine statistical power in a hypothesis test?

level: middleimportance: must knowfreq 70%

basics

~20 s

Power rises with sample size, with the size of the true effect, with a larger alpha, and with lower outcome variance. Three of those are design choices; the true effect is not yours to set, only to assume honestly.

open as a page

A drug lowers systolic blood pressure by 0.4 mmHg with p < 0.001 in 120,000 patients - is that worth adopting?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Almost certainly not. A 0.4 mmHg reduction is far below any clinically meaningful change; the 120,000 patients merely make so tiny an effect distinguishable from zero. The decision needs magnitude weighed against cost and harm, not a test verdict.

open as a page

Why report an effect size alongside the p-value from a hypothesis test?

level: juniorimportance: should knowfreq 66%

basics

~20 s

A p-value only says whether an effect is distinguishable from zero; the effect size says how big it is. Decisions, cost-benefit comparisons and pooling results across studies all need the magnitude, not just the verdict.

open as a page

Why is the odds ratio 2.25 when the relative risk is 1.5 for an event rate rising from 40% to 60%?

level: middleimportance: should knowfreq 43%

basics

~20 s

Odds and probabilities diverge once events are common. Risks of 40% and 60% give a relative risk of 1.5, but the matching odds of 0.667 and 1.5 give an odds ratio of 2.25. An odds ratio always sits further from 1.

open as a page

An early-stage drug safety screen is run at alpha = 0.10 rather than 0.05. How would you defend that choice?

level: seniorimportance: should knowfreq 46%

basics

~20 s

The two errors have very different costs here. The null is that the compound is safe, so a miss advances a possibly harmful compound while a false alarm only triggers extra testing; the looser threshold buys a lower miss rate.

open as a page

Why is post-hoc power computed from an observed effect uninformative?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Observed power is computed from the effect the study just estimated, so for a fixed test it is a one-to-one function of the p-value and adds no new information. A null result always yields low observed power.

open as a page

Why do significant results from underpowered studies overstate the true effect?

level: seniorimportance: should knowfreq 45%

basics

~20 s

In a low-power study only unusually large estimates clear the significance threshold, so the ones that do come from the extreme tail of the sampling distribution. Averaged over such studies, the surviving estimates exaggerate the true effect, sometimes several-fold.

open as a page

What does an eta-squared of 0.06 tell you about a factor in a one-way ANOVA?

level: middleimportance: nice to knowfreq 27%

basics

~20 s

That the factor accounts for 6% of the total variation in the outcome: its sum of squares over the total sum of squares. By Cohen's rough benchmarks that is a medium effect, and 94% of the variation is left unexplained.

open as a page

How would you set an organisation's default significance level when error costs differ sharply by decision?

level: principalimportance: nice to knowfreq 31%

basics

~20 s

Keep one published default for consistency, then allow documented exceptions by decision class, each justified in advance by the cost and recoverability of a false alarm versus a miss. The policy's real job is preventing thresholds chosen after the result is known.

open as a page

Your analysis can reach only 40% power; how do you decide whether to run it at all?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Decide by what the result will be used for. A 40%-power run is defensible as a pilot or as one input to a pooled estimate, and indefensible if a null will be read as no effect.

open as a page