skip to content

When should a hypothesis test use a one-tailed alternative rather than a two-tailed one?

level: middleimportance: must knowfreq 66%

answer

  1. ask what you would do in each direction
  2. the budget is split or concentrated
  3. 1.96 against 1.645 at the same level
  4. the other tail becomes undetectable

basics

~20 s

Use a one-tailed alternative only when a result in the opposite direction would be acted on exactly like no effect, and only when the direction is fixed before the data. Otherwise two-tailed is the honest default.

solid answer

~50 s

Choose the tails from what you would do about a result in each direction, and choose before seeing data. For a coin suspected of favouring heads, a non-directional alternative `H1: p != 0.5` splits the significance level across both tails: at `alpha = 0.05` each tail holds 0.025 and the rejection region is `|z| > 1.96`. A directional alternative `H1: p > 0.5` puts the whole 0.05 in the upper tail, so the threshold drops to `z > 1.645` and the test is more sensitive to an upward effect. The price is total blindness the other way — however strongly the coin favours tails, that test can never declare it. One-tailed is defensible when a move in the other direction would be handled exactly as no move at all; when both directions would trigger action, use two tails.

go deeper

for a junior

Know that two-tailed is the default and that the alternative can be written with != , > or <. Be able to say which form matches a scenario when the interviewer describes it.

for a middle

Explain the probability budget: the same alpha either splits into two tails or concentrates in one, which is why the thresholds are 1.96 and 1.645, and state what the one-tailed test gives up.

for a senior

Justify the choice from the decision it feeds — what action follows a result in each direction — and be ready to push back on a one-tailed design proposed on the grounds of expectation.

for a principal

Set the team convention: two-tailed by default, one-tailed only with a written rationale agreed before data collection, so the choice never becomes a lever people reach for when a result is borderline.

## What the tails actually are The significance level `alpha` is a budget of probability: the share of outcomes, computed assuming the null is true, that you agree in advance to call "surprising enough to reject". The **rejection region** is where you spend that budget on the scale of the test statistic. One- versus two-tailed is simply the decision of whether to spend it all at one end or split it between both. Take a coin suspected of being unfair, with `H0: p = 0.5`, and a standardised test statistic `z`. - **Non-directional (two-tailed):** `H1: p != 0.5`. The 0.05 budget splits into 0.025 in each tail. The rejection region is `|z| > 1.96` — reject if the coin looks strongly biased in *either* direction. - **Directional (one-tailed):** `H1: p > 0.5`. The entire 0.05 goes into the upper tail. The rejection region is `z > 1.645` — reject only for strong evidence of heads-favouring bias. Notice that 1.645 is smaller than 1.96. Concentrating the budget in one tail lets you reject on weaker evidence *in that direction*, because you gave up the right to reject in the other one. ## The tradeoff, stated honestly The one-tailed test is more sensitive to a true effect in the chosen direction: the bar it must clear moves from 1.96 to 1.645, roughly a 16% shorter standardised distance. That is a real but modest gain — far smaller than what a larger sample buys. What it costs is absolute. Under `H1: p > 0.5`, a coin that comes up tails 90 times in 100 produces a large negative `z`, which is not in the rejection region at all. The test cannot report that result as significant, no matter how extreme it is. You have not merely deprioritised the other direction; you have made it invisible. So the decision rule is behavioural, not statistical: **would a result in the opposite direction change what you do?** If yes — if underfilling and overfilling both trigger an intervention, if a metric moving down would be as newsworthy as it moving up — the alternative has to cover both directions. If a downward result would be handled identically to no result, a one-tailed alternative honestly reflects the decision problem. In practice that condition is rarely met, which is why two-tailed is the default in most fields, and why a one-tailed test in a report invites the question of why. ## Writing the pair correctly The two hypotheses must partition the parameter space: - Two-tailed: `H0: p = 0.5` against `H1: p != 0.5`. - One-tailed upward: `H0: p <= 0.5` against `H1: p > 0.5`. The test is executed at the boundary `p = 0.5`, the value inside the null that is hardest to reject. - One-tailed downward: `H0: p >= 0.5` against `H1: p < 0.5`, rejecting when `z < -1.645` at `alpha = 0.05`. A frequent slip is to write `H1: p > 0.5` while still comparing against the two-tailed critical value, or the reverse — quoting 1.96 for a one-sided test. The alternative and the critical value are two halves of the same design choice and must agree. ## Where the critical values come from For a standard normal reference distribution, 1.645 is the point with 5% of the probability above it and 1.96 is the point with 2.5% above it — hence 5% split across `|z| > 1.96`. Tightening to `alpha = 0.01` moves the one-sided threshold to about 2.33 and the two-sided threshold to about 2.58. Loosening to `alpha = 0.10` moves the one-sided threshold to about 1.28. The pattern to remember: **for the same `alpha`, the one-sided critical value is always closer to zero than the two-sided one**, because it is not sharing the budget. ## Order of operations Both `alpha` and the number of tails belong to the design phase, alongside the hypotheses themselves. They are choices about what evidence would persuade you, and that question is only meaningful before you know what the evidence says. Deciding tails after inspecting the data does not merely look bad — it changes the actual error rate of the procedure while leaving the reported one unchanged. A clean write-up therefore states, before any results: the null, the alternative including its direction, the significance level, and the resulting critical value. Everything after that is arithmetic. ## Interview shorthand If asked to justify a choice on the spot: name the two hypotheses, say which direction would prompt action, and only then pick the tails. Candidates who lead with "one-tailed, because we expect it to go up" are giving an expectation as a justification; expectation is exactly the wrong criterion, because a surprise in the other direction is often the most important thing the data could tell you.

  • How much extra sensitivity does a one-tailed test actually buy?
    Modest. At `alpha = 0.05` the threshold moves from 1.96 to 1.645, about a 16% shorter standardised distance to clear, which translates into a small gain in detecting a true effect in the chosen direction. It is never a substitute for collecting more data, and it costs the entire ability to detect an effect the other way.
  • A stakeholder says a drop in the metric would be catastrophic. Does that argue for a one-tailed test?
    The opposite. If a drop matters, you must be able to detect it, so the alternative has to cover both directions. A one-tailed alternative is defensible only when a move the other way would be handled identically to no move at all — which plainly is not the case once someone tells you the downside is catastrophic.
  • How do you write a one-sided null and where is the test actually carried out?
    Write `H0: p <= 0.5` against `H1: p > 0.5`. The test is performed at the boundary, `p = 0.5`, because that is the value inside the null that is hardest to reject. If the evidence is strong enough to reject at the boundary, it is stronger still against every value further inside the null region, so the boundary controls the error rate for the whole set.

saying these in an interview costs you the question

  • Picks one-tailed because it makes significance easier to reach.
  • Thinks a two-tailed test puts the full alpha in each tail.
  • Believes a one-tailed test can still flag the opposite direction.
  • Quotes 1.96 as the threshold for a one-sided test at 0.05.
  • Justifies the direction by what the team expects to happen.

context