How should a launch committee set interval-based ship rules for low-traffic A/B tests?
answer
- precision scales with square root of traffic
- decide before running whether to run
- cost-weighted, asymmetric bars
- reframe as ruling out harm
- never log uninformative as null
basics
~20 sDecide up front what a wide interval licenses. On low-traffic surfaces, pre-commit to a rule such as ship if the interval excludes meaningful harm and the change is cheap, and never log an uninformative test as evidence of no effect.
solid answer
~50 sLow-traffic surfaces routinely cannot deliver an interval narrower than the lift anyone cares about, so a committee that demands significance there simply never ships. The workable policy has three parts. First, decide before launch whether the surface can achieve the precision the decision needs; if it cannot, say so and choose a different evidence standard rather than running a test whose readout is predetermined to be ambiguous. Second, define what a wide interval licenses — commonly a guardrail rule: ship a low-cost change when the interval excludes harm beyond an agreed bound, defer an expensive or risky one regardless of the point estimate. Third, forbid rewriting an uninformative readout as a null result; log the achieved precision so nobody later cites it as proof the idea failed. The alternative to statistical evidence is explicit judgment, not a test dressed up as one.
go deeper
Be ready to explain why a low-traffic experiment produces a wide interval and why that width, rather than the point estimate, decides what the result can support.
Describe the arithmetic that links traffic to interval width, and explain why a fixed significance bar applied everywhere makes small surfaces untestable in practice.
Show that you check achievable precision against the decision threshold at planning time, and can propose a defensible alternative — pooling, a proxy metric, or a harm-only safety check — when the surface cannot support the test.
Own the policy: cost-tiered, pre-committed rules for what a wide interval licenses, a record of achieved precision on every readout, and a vocabulary that keeps inconclusive results distinct from genuine null findings.
## The structural problem The precision of a lift estimate improves roughly with the square root of the sample size. A surface with a hundredth of the traffic of the main funnel therefore delivers an interval about ten times wider for the same test duration. On such surfaces an interval of plus or minus 8% on a metric where a 2% lift would be excellent is normal, not a failure of execution. A committee that applies one rule — ship when the interval excludes zero — gets two bad behaviours. Low-traffic teams either never ship, or they run tests long past any sensible horizon, peeking until an interval finally clears zero, which inflates the wrong-ship rate far beyond the nominal level. Both outcomes are worse than an explicit policy. ## Deciding before the test whether to test The first policy question is answered at planning time, not readout: given this surface's traffic and this metric's variance, what interval half-width is achievable in an acceptable window? Hold that against the lift the decision turns on. - If achievable precision is comfortably finer than the decision threshold, run a normal experiment. - If it is coarser, the readout is predetermined to be ambiguous. Running the test anyway buys a number that will be argued over. Better options: pool related surfaces into one test if the change is the same, move the readout to a higher-traffic proxy metric that is genuinely upstream of the outcome, extend the horizon deliberately with a pre-declared end date, or accept that this decision will be made on judgment and run the test only as a safety check. Making this a required field in the test plan — achievable half-width versus decision threshold — is the single highest-leverage rule available, because it moves the argument to before anyone is invested in a result. ## What a wide interval licenses When a test that could not have been precise comes back wide, the committee needs a rule that was written down earlier. Two shapes work. **Asymmetric, cost-weighted shipping.** The bar depends on what a wrong decision costs. For a low-cost, easily reversible change on a small surface, requiring proof of a win is disproportionate; requiring that the interval exclude harm beyond an agreed bound is enough, and the point estimate being positive is a tiebreaker rather than evidence. For an expensive or hard-to-reverse change, a wide interval means the evidence standard was not met and the decision reverts to explicit judgment with the uncertainty on the record. **A safety-check framing.** Reframe what the experiment is for. Instead of asking it to prove a win it cannot resolve, ask it to rule out a regression larger than an agreed bound. That question needs less precision, since the bound is usually larger than the lift being chased, and a test that answers it is genuinely informative even when it says nothing about upside. Both rules must be fixed before the readout. A threshold chosen after the interval is on screen is not a rule. ## Institutional hygiene Three practices keep interval-based decisions honest over years rather than weeks. - **Log achieved precision, always.** Every readout records the interval and the half-width it reached. Without this, an uninformative test becomes "we tried that, it did nothing" and the idea is dead for the wrong reason. - **Separate the words.** Distinguish *conclusively no meaningful effect* — interval entirely inside the band nobody would act on — from *inconclusive* — interval spanning the decision threshold. Committees that use one phrase for both slowly lose the ability to tell which experiments taught them anything. - **Cap the extend-until-it-works reflex.** Either declare the end condition at launch, or restart the test with an honest sample-size plan. Repeated look-and-extend is where nominal error rates and real error rates diverge most. ## The tradeoff to own Relaxing the bar for low-traffic surfaces means shipping some changes that do nothing, and occasionally one that mildly hurts. Holding the bar means shipping nothing there and letting the surface stagnate. Neither is free, and the choice is the committee's to make explicitly, per cost tier, in advance. What is not defensible is leaving the rule unstated and settling each case by whichever argument is made most forcefully in the meeting — because that reliably converges on shipping whatever its sponsor wanted, with an interval quoted as decoration.
- How do you stop teams from extending a low-traffic test until the interval finally excludes zero?Require an end condition in the test plan — a fixed sample size or date — and treat any extension as a restart that must be declared. Repeated look-and-extend drives the real wrong-ship rate well above the nominal level, and the only durable defence is that the stopping point was fixed before the data were seen.
- Is pooling several low-traffic surfaces into one test a sound way to buy precision?It is, when the change is genuinely the same intervention and the surfaces are similar enough that one pooled lift is a meaningful quantity. The cost is that a heterogeneous true effect gets averaged away, so the pooled estimate can hide a win on one surface and a loss on another. Declare the pooling before launch.
- What should the committee record when it ships a change on judgment rather than evidence?That it did exactly that: the interval achieved, the threshold it failed to resolve, and the non-statistical reasons for shipping. This keeps the decision reviewable and stops the experiment from being cited later as the evidence base for something it never established.
saying these in an interview costs you the question
- Applying one significance bar to every surface
- Choosing the decision threshold after seeing the interval
- Recording an uninformative test as a null result
- Extending a test until the interval excludes zero
- Treating a wide interval's point estimate as the answer