skip to content

Which of two competing claims should be the null hypothesis, and how do you decide?

level: principalimportance: nice to knowfreq 33%

answer

  1. ask who has to convince whom
  2. one side is protected by the guarantee
  3. the null must be computable
  4. the claim to be earned goes in H1

basics

~20 s

Put the status-quo, no-effect claim in the null and the claim that must earn its conclusion in the alternative. The null is the protected side and must name a specific value, so placement decides the burden of proof.

solid answer

~50 s

The convention is that the null holds the status quo, the no-effect claim, or whatever costs nothing to keep — but the real decision is about burden of proof. The framework protects whatever sits in `H0`: it bounds how often a true null is wrongly rejected, and offers no symmetric bound the other way. So the claim you want someone to have to earn belongs in `H1`. There is also a hard constraint: `H0` must pin down a specific value so a distribution exists to test against — "there is some effect of unknown size" cannot be a null. When the claim you must demonstrate *is* the no-difference one, the standard setup cannot deliver it by failing to reject; you need a design where that claim is what the data must support against a threshold stated in advance.

go deeper

for a junior

Learn the default and apply it: the null is the no-effect or status-quo claim, the alternative is what you would need evidence to support. Recognising the pattern in a scenario is enough at this stage.

for a middle

Explain why the default exists — the null has to name a value precise enough to compute under, which is why the vague claim ends up on the alternative side.

for a senior

Recognise when the conventional framing works against the question being asked, especially when a stakeholder needs sameness demonstrated rather than difference detected, and say what design you would use instead.

for a principal

Own the burden-of-proof policy across decision types and require hypotheses, level and tails to be fixed in the plan, so framing never becomes a per-project lever on the outcome.

## Two rules, one convention Most textbooks give the convention — the null is "no effect", "no difference", the status quo — and stop there. The convention is right most of the time, but the reasoning underneath it is what an interviewer at this level is probing, because it is what tells you when the convention does not apply. There are two genuine constraints. **Constraint 1: the null must be computable.** The test measures how surprising the data would be *if the null held*, which requires the null to determine a distribution. `mu = 500` does. `p = 0.5` does. "There is an effect, of some unknown size" does not — it is a family of infinitely many distributions with no single reference point. This alone forces the sharp, no-effect-shaped claim into `H0` in most testing problems, and it is why you cannot simply swap the labels to reverse the burden of proof. **Constraint 2: the null is the protected side.** The procedure guarantees that a true null is wrongly abandoned at most `alpha` of the time — 5%, by common convention. There is no matching guarantee that a false null gets caught; whether it does depends on how large the real effect is and how much data you gathered. The null therefore enjoys a protection the alternative does not. Whatever you place there will be kept unless the evidence is strong. Put together: **the claim that should be hard to establish goes in the alternative; the claim that should survive unless the evidence is strong goes in the null.** ## Reading the choice as burden of proof Once framed this way, the question becomes organisational rather than mathematical. Who should have to convince whom? On a bottling line advertised at 500 ml, the natural setup is `H0: mu = 500` against `H1: mu != 500`. The consequence is that the plant keeps running unless quality control produces strong evidence of drift. That is a deliberate stance: stopping a line is expensive, so the burden is placed on the challenger, and the price is that a small persistent drift may go unflagged for a long time. Invert the situation and the stance becomes uncomfortable. Suppose a regulator's question is "demonstrate that this line is correctly calibrated". Placing "the line is calibrated" in the null means the line passes whenever the evidence is weak — including when the sample was tiny or the measurement sloppy. The worst-designed study produces the most reassuring verdict. That is the wrong incentive, and it is a direct consequence of putting the claim-to-be-demonstrated on the protected side. ## When you must demonstrate sameness This is the case that separates strong answers from recited conventions. If the conclusion you owe someone is "no meaningful difference", the ordinary setup cannot supply it, because failing to reject is what an uninformative study does by default. The way out is not to relabel the hypotheses — constraint 1 forbids putting a vague "there is a difference" in the null. It is to **restate the claim against a concrete threshold**: decide, before data collection and on substantive grounds, how large a difference would actually matter, and build the test so that the data must positively support the difference being smaller than that. The threshold is the part that requires judgment and the part that must not be chosen after seeing results. The structural point for an interview is that this is a different design, committed to in advance, and never a reinterpretation of a non-significant result. ## What a lead is expected to own A few decisions in this area belong to whoever sets the standard, not to the analyst running the numbers. - **Direction of the burden per class of decision.** Launch decisions, safety checks and regulatory claims do not all deserve the same default. Deciding which of them place the burden on the challenger, and which require the change itself to be affirmatively supported, is a policy question you should be able to argue. - **Consistency.** If the burden is arranged case by case, whoever writes the analysis plan effectively decides the outcome. A standard that says which claim is the null for each recurring decision type removes that discretion. - **Timing.** Hypotheses, level and tails belong in the plan before data collection. Any of them chosen afterwards silently changes the error rate of the procedure while leaving the reported figure intact. - **Honesty about what non-rejection buys.** A team that treats surviving nulls as established claims will accumulate false reassurance in proportion to how weak its studies are. ## The short version Ask what the framework should protect. The protected claim, which must also be specific enough to compute under, becomes `H0`. The claim that must be earned becomes `H1`. If the claim you are required to demonstrate turns out to be the protected one, you do not have a labelling problem — you have the wrong design, and you need a threshold-based one instead.

  • Why can't you simply swap H0 and H1 to shift the burden of proof?
    Because the null must specify a distribution and the alternative need not. "The effect is exactly zero" is computable; "the effect is non-zero, size unknown" is not, so it cannot serve as a null. Shifting the burden requires restating the claim in terms of a concrete threshold that *can* sit in a null, not a bare reversal of labels.
  • What does it mean that the framework protects whatever sits in the null?
    The procedure bounds how often a true null is wrongly rejected, at the significance level you chose. There is no equally small bound on failing to reject a false null — that depends on the true effect size and the amount of data. So the null side enjoys a guarantee the alternative never gets, which is why the placement is a design decision rather than a formality.
  • How do you keep hypothesis framing consistent across a team?
    Standardise it per decision type rather than per project: name, for each recurring decision, which claim is the null and which direction carries the burden, and require hypotheses, level and tails to be written in the analysis plan before data collection. Case-by-case framing hands whoever writes the plan an unaccountable lever over the outcome.

saying these in an interview costs you the question

  • Puts the claim to be demonstrated in the null hypothesis.
  • Assumes the H0 and H1 labels are freely interchangeable.
  • States a null with no specific value to compute under.
  • Thinks a surviving null establishes the claim it makes.
  • Chooses which claim is the null after seeing the results.

context