On day 2 of a planned 14-day A/B test the dashboard shows p = 0.04 — do you ship?
answer
- the threshold assumes one look
- you planned 14 days for a reason
- day-2 direction is not a decision
- harm guardrails are a separate rule
- record it, wait, decide on the date
basics
~20 sNo. The 0.05 threshold is calibrated for a single analysis at the planned sample size, so a day-2 reading is not the test you designed. Run to the agreed horizon unless the plan already allowed an early call.
solid answer
~50 sI would not ship. The test was planned for 14 days, which means the significance threshold is calibrated for one analysis at that horizon; reading it on day 2 and acting is optional stopping, and it turns a nominal 5% false-positive rate into something several times larger. A day-2 sample is also small, so even if the effect is real the estimate is unstable — the same data stream can read as a win now and as flat or negative a week later. What I would say to the team is: the current direction is encouraging, here is how wide the interval still is, and the answer arrives on the planned date. The exception is a pre-agreed one — a harm guardrail that fires on a bad result, or a design chosen before launch that supports early looks.
go deeper
Recall the rule and the reason: the 0.05 bar assumes one analysis at the planned sample size, so an early reading is not that test. Say you would wait for the agreed end date.
Explain what breaks mechanically — the stopping time became data-dependent — and quantify it. Also note that a small day-2 sample makes the effect size unreliable even when the direction is real.
Show you can hold the line usefully: give the stakeholder the direction, the interval width and the decision date, and separate a pre-agreed harm guardrail from an improvised early win call.
Frame it as policy rather than a single argument. Decide who owns the stopping rule, what standing exception exists for harm, and whether the org's need for speed justifies buying a design that supports early decisions.
## The question behind the question An interviewer asking this is not testing whether you can read a p-value. They are testing whether you understand that **a threshold is only meaningful together with the protocol that produced it**, and whether you will hold that line when the number is exciting. ## Why day-2 significance is not evidence of a win When the test was planned for 14 days, an implicit contract was signed: *we will collect roughly this much data and analyse once at the end.* Under that contract, and assuming the two arms are truly identical, the procedure declares a winner about 5% of the time. That is what "p < 0.05" buys you. Reading the dashboard on day 2 and acting changes the procedure into: *analyse repeatedly, stop the first time it crosses.* The 5% guarantee does not survive the change. The statistic on a growing sample is a noisy trajectory; a decision rule that fires on its first excursion catches noise far more often than a single scheduled reading does. Checking daily through a two-week test and shipping at the first crossing runs at roughly a one-in-five false-positive rate rather than one in twenty. There is a second, independent problem. On day 2 the sample is small, so the confidence interval is wide. Even in the happy case where the effect is genuinely positive, the *magnitude* you would quote to the business is unreliable, and "we saw +8%" becomes a number people plan headcount around. ## What a strong answer sounds like A weak answer is "no, that's peeking" and nothing else — technically correct, professionally useless, because the person asking has a decision to make. A strong answer refuses the shipping decision while still being helpful: 1. **State the position.** We do not call the test on day 2; the result we designed arrives on day 14. 2. **Say why in one sentence, without jargon.** The 5% bar assumes we look once at the end. Looking every day and acting on the first good number means we would be wrong far more often than 5% of the time. 3. **Give them what is honestly available now.** The direction so far, how wide the interval still is, and whether anything looks broken. Direction is information; a shipping decision is not. 4. **Name the date.** "We will have the answer on the 14th" converts an argument into a wait. 5. **Name the legitimate exceptions.** Guardrails that detect harm are a different decision with a different cost profile — if the variant is hurting users you stop it, and you are not claiming a calibrated win when you do. And if the business genuinely cannot wait two weeks for this class of decision, that is a design requirement to raise *before* the next launch, when a monitoring-capable method can be chosen deliberately. ## The asymmetry between stopping for a win and stopping for harm This distinction is worth having ready, because it is the most common follow-up. Stopping early **for a win** is a claim: *this variant is better, ship it to everyone.* That claim leans on a calibrated error rate you have just invalidated. Stopping early **for harm** is a different act: *this looks bad enough that continuing costs more than the information is worth.* You are not asserting statistical proof of harm; you are making a cost decision under uncertainty, and you accept that some of those stops will turn out to have been unnecessary. The false-positive inflation still exists, but the loss function is asymmetric — shipping a fake win is expensive and hard to undo, while pausing a variant that was actually fine costs you a rerun. Say this explicitly rather than pretending the two cases are symmetric. ## The trap to avoid The seductive middle position is "let's keep it running and check again tomorrow — if it's still significant, we ship." This feels like caution and is not. Requiring two crossings in a row is still a data-dependent stopping rule, and nobody has computed its error rate. If the answer is going to be decided by looking, it needs a boundary designed for looking, not a personal sense of how many consecutive green days feel convincing. The other trap is treating an early non-significant result as licence to extend the run "until it turns." That is the same defect wearing different clothes: the data is choosing the sample size. Either way, the cure is that the horizon and the decision rule were written down before launch and are honoured afterwards. ## What you would actually do Let it run. Record the day-2 reading in the test log so nobody later claims the result was known early. Check that the experiment is healthy — instrumentation firing, arms filling as expected — because *operational* problems genuinely are worth catching on day 2, and that is a legitimate reason to look at an experiment early. Then decide on the date you said you would.
- What if the day-2 result is a large drop instead of a gain — do you still wait?No, and that is not inconsistent. Stopping for harm is a cost decision, not a statistical claim: continuing to expose users to a variant that looks damaging costs more than the information is worth, and you accept that some of those stops were unnecessary. Ideally the harm guardrail and its trigger were written into the plan before launch so the call is not improvised under pressure.
- The stakeholder proposes waiting one more day and shipping if it is still significant. Is that a reasonable compromise?No. Requiring two consecutive crossings is still a data-dependent stopping rule, just an undocumented one whose error rate nobody has computed. It feels cautious because it demands more evidence, but it was invented after seeing the data. If early decisions are genuinely needed, the boundary has to be designed for repeated looks before the test starts.
- Is there anything you legitimately should be checking on day 2?Yes — the health of the experiment rather than its outcome. Is instrumentation firing, are both arms receiving traffic, are there errors concentrated in one variant. Catching a broken setup on day 2 saves the whole run, and none of it involves acting on the effect estimate. Looking at operational diagnostics is not what peeking means.
saying these in an interview costs you the question
- Ships because the p-value already cleared the bar
- Says two weeks was only a rough guess, so day 2 counts
- Proposes waiting for a second significant day as a compromise
- Cannot distinguish stopping for harm from stopping for a win
- Refuses with jargon and offers the stakeholder nothing