An analyst switches to a one-tailed test after seeing the data trend upward — what is wrong?
answer
- the data chose the decision rule
- the other branch still counts
- the region is really both directions
- 0.05 in each direction totals 0.10
basics
~20 sThe reported significance level is no longer the one the procedure delivers. Letting the data choose the tail makes the effective rejection region both tails at the one-sided threshold, so a stated 5% false-positive rate is really about 10%.
solid answer
~50 sThe stated `alpha` stops being true. Fixing `H1: mu > mu0` in advance spends the whole 5% budget in one tail, with a critical value of 1.645. Letting the data pick which tail to use means the analyst would have rejected a large positive `z` *or* a large negative one — so the real rejection region is `|z| > 1.645`, which under the null has probability about 0.10. The error rate silently doubles while the report still claims 5%. Because 1.645 sits below the two-tailed 1.96, borderline results that were not significant now cross the line, which is exactly the class of result the switch tends to be applied to. The remedy is to fix the tails and the level before collecting data and record that choice; where it was not done, report the two-tailed result.
go deeper
Know the rule as a rule: the number of tails and the significance level are chosen before the data, and changing them after looking is not allowed.
Be able to compute the damage — both branches of the switch are live under the null, so the effective region is both tails at 1.645 and the real error rate is about 0.10.
Show how you would find this in review and what you would do about it: ask when the direction was fixed, check whether the statistic sits in the 1.645-to-1.96 band, and report the two-tailed result.
Own the process fix rather than the individual catch — a template where hypotheses, tails and level are recorded before data collection removes the temptation and makes the omission visible.
## What the switch really does A significance level is a property of a **procedure**, not of a number in a report. The procedure here is: "look at the data, see which direction the effect went, then run a one-tailed test in that direction at 0.05". Ask what that procedure does when the null is true. If the sample happens to land high, the analyst rejects when `z > 1.645`, which has probability 0.05. If it lands low, the analyst rejects when `z < -1.645`, also probability 0.05. The two cases are disjoint and both are live before the data arrive, so the total probability of rejecting a true null is about **0.10**, not 0.05. The rejection region of the procedure actually is `|z| > 1.645`. That is a two-tailed test run at a significance level of 0.10 while being reported as a one-tailed test at 0.05. ## Why the damage lands exactly where it hurts The inflation is not spread evenly across possible results. Compare thresholds at a nominal 0.05: - Two-tailed, chosen in advance: reject when `|z| > 1.96`. - Tail chosen after the fact: reject when `|z| > 1.645`. Every result with `|z|` between 1.645 and 1.96 flips from non-significant to significant. That band is precisely where marginal, ambiguous findings live — the ones a team is most tempted to rescue, and the ones least likely to replicate. The switch does not slightly relax a strict rule; it converts a specific population of borderline results into claimed discoveries. ## Why the intuition feels innocent The analyst's reasoning usually sounds like this: "we always expected the metric to go up, the data confirm it, so testing the other direction was never meaningful — the one-tailed test just reflects what we knew". The flaw is that the claim "we always expected up" is unverifiable after the fact and, crucially, is not what the procedure did. Had the data gone down, the same analyst would have had a story for that too — and that counterfactual branch is part of the error rate whether or not it was taken. Error rates are computed over the procedure's behaviour across all datasets it might have seen, not over the one that arrived. This is also why "but the direction really was obvious" is not a defence. If it was obvious, it could have been written down beforehand at no cost, and then the 0.05 claim would be true. ## Doing it properly The tails, the significance level and the resulting critical value belong to the design phase, next to the hypotheses. A defensible analysis plan says, before data collection: `H0: mu = mu0`; `H1: mu > mu0`; `alpha = 0.05`; reject when `z > 1.645`. Written that way, the one-tailed test is entirely legitimate, and the reported error rate is the real one. The substantive condition for choosing one tail has not changed either: a result in the excluded direction must be one you would treat exactly as no effect. If a downward move would trigger a rollback, an investigation or a headline, the alternative had to be two-sided from the start. ## What to do when it has already happened There is no post-hoc correction that restores the original claim, because the information about when the choice was made is gone. Practical options: - **Report the two-tailed result** at the level fixed in advance. This is honest and usually enough; if the finding survives at `|z| > 1.96`, nothing was lost. - **Report both**, stating plainly that the direction was selected after inspecting the data, so the one-tailed figure carries an effective error rate of roughly 0.10. - **Treat the finding as exploratory** and confirm it on fresh data with the direction fixed in advance. This is the only route that recovers a genuine 0.05 claim. What not to do: quietly present the one-tailed number, or argue the correction is unnecessary because the effect was large. If the effect was large it clears the two-tailed threshold anyway, and the argument is unnecessary. ## Spotting it in review The signature is a one-tailed test whose direction happens to match the observed effect, with no design document that fixed the direction. Useful review questions: - Where is the direction of the alternative stated, and does it appear anywhere before the results? - Would a result in the opposite direction have been reported at all, or explained away? - Is the test statistic in the 1.645-to-1.96 band — that is, does the tail choice decide the verdict? A report that discusses the tail choice only in its results section, and never in its design section, has effectively answered the first question. ## The general principle Anything that lets the observed data influence the decision rule inflates the false-positive rate of that rule, even when each individual step looks reasonable in isolation. The tail switch is the cleanest illustration because the inflation is exactly computable: 0.05 becomes 0.10, and the report never mentions it.
- Is a one-tailed test ever legitimate, then?Yes, when the direction is fixed and written down before any data are seen, and when a result in the opposite direction would genuinely be treated as no effect. The problem is never the one-tailed test itself — it is choosing the tail with knowledge of the data, which turns a stated 5% error rate into an actual 10% while the report still says 5%.
- How would you detect this in a colleague's analysis?Look for a one-tailed test whose direction matches the observed effect, with no record of the direction being set in advance. Ask when the tails and the level were decided and where that is documented. A write-up that mentions the tail choice only in the results, never in the design, is the usual signature — especially when the statistic sits between 1.645 and 1.96.
- The analyst argues the effect is so large the tail choice cannot matter. Is that a defence?It is self-defeating. If the statistic clears the two-tailed threshold of 1.96, then reporting the two-tailed result costs nothing and the argument is unnecessary. The argument only ever gets made when the result sits in the 1.645-to-1.96 band — exactly the range where the tail choice, not the evidence, is deciding the verdict.
saying these in an interview costs you the question
- Claims the stated 5% level still holds after switching tails.
- Defends the switch by pointing at the observed effect direction.
- Says the direction was always expected, with nothing written down.
- Argues no correction is needed because the effect looks large.
- Presents only the one-tailed figure without disclosing the switch.