skip to content

An executive asks 'is it significant yet?' on day 2 of a planned 14-day test — how do you handle it?

level: principalimportance: should knowfreq 36%

answer

  1. answer the decision, not the number
  2. share direction, withhold the verdict
  3. price the request in business terms
  4. much peeking is fear of harm
  5. fix it at planning, not in the argument

basics

~20 s

Answer the decision behind the question, not the number. Say what is honestly known now, give the date the result arrives, and treat recurring demands for early reads as a planning problem to fix before the next launch.

solid answer

~50 s

I would not answer with the number and I would not answer with a lecture. First, find the decision: someone usually asks this because a launch date, a budget or a competitor is forcing their hand, and that is the thing to address. Second, give what honestly exists — the direction so far, how wide the interval still is, whether anything looks broken, and the date the result lands. Third, be explicit about the cost of what is being asked: acting on the first good day means being wrong perhaps one time in five rather than one in twenty, and those wrong calls are expensive to unwind. Fourth, if the underlying constraint is real and recurring, fix it upstream — smaller-scope decisions, an agreed harm-stop policy, or committing before launch to a design that supports early reads. Winning this argument once and rerunning it every quarter is a failure.

go deeper

for a junior

Know not to hand over the interim number, and to reply with the decision date instead. Escalate the pressure to whoever owns the test plan rather than absorbing it alone.

for a middle

Be able to state the cost of an early call in plain terms — roughly one wrong ship in five instead of one in twenty — and to offer direction and uncertainty width without offering a verdict.

for a senior

Show you can find the decision behind the question, distinguish a harm concern from impatience, and keep the stakeholder informed enough that they do not go around you for the numbers.

for a principal

Own the systemic fix: pre-commitment agreed with stakeholders before launch, a standing harm policy, published base rates on how often early calls held up, and a deliberate decision about whether faster experimentation is worth buying.

## Why this is a leadership question, not a statistics question The statistics here take one sentence: reading the result early and acting on it turns a 5% error rate into something several times larger. Anyone senior enough to be asked this already knows that. What is being tested is whether you can protect the integrity of the decision process **without becoming the person who says no and offers nothing**, and whether you can fix the underlying condition rather than relitigating it every fortnight. An analyst who answers "that would be peeking" and stops has technically defended the method and practically lost it — the next test will be called early by someone who did not ask. ## Step one: find the decision under the question Nobody wants a p-value. They want to know whether to hold a launch, release headcount, brief a partner, or tell a board something. Ask directly: *what decision are you trying to make, and when do you have to make it?* This usually resolves into one of a few cases, and they need different answers: - **No real deadline.** The question is habit or anxiety. The answer is the decision date plus a promise to bring it unprompted. - **A real deadline before the horizon.** Now you have a genuine conflict, and the honest framing is a tradeoff: we can decide on the 5th with substantially worse odds of being right, or on the 14th with the odds we designed for. Put numbers on both and let the decision-maker own it explicitly. - **A decision that does not actually need this test.** Often the launch can proceed on a reversible basis, or the choice hinges on something the experiment was never measuring. Then the conflict evaporates. - **Something looks wrong.** If the concern is that the variant is hurting users, that is a harm question with its own rule, and it should be answered immediately and separately from any claim of a win. ## Step two: say what is honestly available Refusing to share anything reads as gatekeeping and invites someone else to pull the numbers. What you can share on day 2 without endorsing a decision: - The observed direction, stated as a direction and not a magnitude. - How wide the uncertainty still is — "the range consistent with our data today spans everything from a meaningful loss to a large win" communicates more than any p-value. - Whether the experiment is healthy and on track to deliver on time. - The exact date the answer arrives, and a commitment to bring it without being chased. ## Step three: price the request rather than refusing it The strongest move is to convert the argument into a cost the business can weigh. Something like: *if we decide on the first good day, we ship things that do nothing roughly one time in five instead of one in twenty, and the number we quote will also be inflated, so the follow-up forecast will miss.* That is a business statement, and executives are equipped to make business tradeoffs. It also leaves the door open: sometimes the answer really is "we accept those odds", and if it is made knowingly, in writing, that is a legitimate call — the failure mode is making it accidentally by looking at a dashboard. ## Step four: fix the recurring cause If this conversation happens every quarter, the process is wrong, not the executive. The durable moves: - **Pre-commit at launch.** The decision date and rule are agreed with stakeholders before traffic starts, so the day-2 conversation refers to a shared commitment rather than to your professional preference. - **Set a standing harm policy.** When everyone knows a bad variant will be stopped immediately, the anxiety that produces the day-2 question drops sharply. Much early peeking is fear, not impatience. - **Buy the capability if speed matters.** If the organisation genuinely needs to act on experiments faster, that is a design decision to make deliberately before launch rather than a discipline problem to keep losing. - **Publish the base rates.** Showing how often early calls have failed to reproduce is more persuasive than any explanation of sampling distributions, and it turns the argument into evidence about your own organisation. ## What a weak answer looks like Giving the number "just as a heads-up, it's at 0.04" — once said, it has been acted on. Refusing in jargon. Promising to check again tomorrow, which invents a stopping rule on the spot. Or agreeing to ship while privately expecting to be vindicated later; the follow-up never happens, and the inflated estimate stays in the forecast. ## What a strong answer looks like Calm, specific and short: here is what we know, here is what we do not, here is the date, here is the cost of deciding sooner, and here is what I will change so we are not having this conversation next quarter. You are protecting a process rather than defending a rule, and you are treating the executive as someone capable of weighing a tradeoff once it is stated plainly.

  • The executive accepts the odds and wants to ship anyway. What do you do?
    Make it an explicit, recorded decision rather than a drift. State the odds in writing, note that the reported effect size is likely inflated so it should not enter a forecast, and propose a follow-up measurement after launch. Then support the call. A knowingly accepted risk owned by the decision-maker is legitimate; the failure mode is the same outcome arrived at by nobody deciding anything.
  • How do you stop this conversation recurring every quarter?
    Move it upstream. Agree the decision date and rule with stakeholders before traffic starts, publish a standing harm-stop policy so people stop watching out of fear, and if the business genuinely needs faster answers, commit to a design that supports early reads before launch. Also publish how often past early calls failed to reproduce — internal base rates persuade where explanations of sampling distributions do not.
  • Is there ever a case where you volunteer the interim numbers?
    For harm, yes — if the variant looks like it is damaging users, that is a separate decision with its own rule and it should be raised immediately rather than waited out. Otherwise no, because once an interim estimate is spoken it has already influenced the decision, whatever caveats came with it. Direction and uncertainty width can be shared without handing over a verdict.

saying these in an interview costs you the question

  • Shares the interim p-value with a verbal caveat attached
  • Refuses using statistical jargon and offers no alternative
  • Promises to check again tomorrow and decide then
  • Treats it as a discipline problem rather than a planning failure
  • Agrees to ship early and never schedules a re-measurement

context