Your experiment hits its planned sample size in three days. Why keep it running to the planned end date?
answer
- count and calendar are two different requirements
- three days is Monday to Wednesday
- early arrivals are not the eventual population
- frequent visitors are exposed almost immediately
- surplus traffic is a capacity signal
basics
~20 sA sample count and a calendar window are different requirements. Three days covers only three weekdays and over-represents your most frequent visitors, so the sample is large but not representative. Read the result when both are met.
solid answer
~50 sHitting the sample target is necessary but not sufficient. Three days is Monday to Wednesday, so the result speaks only for mid-week behaviour and says nothing about weekend users. There is a second, less obvious problem: within any window, the users who accrue first are your heaviest users, because frequent visitors have more chances to be exposed early. Light and lapsed users arrive later and behave differently, so a three-day sample is skewed towards a population the rollout will not be limited to. Both effects push the estimate away from the full-week, full-population number you actually want. The right framing is that the plan carries two preconditions — the sample target and the calendar end date — and you read the result when both are satisfied. If the sample arrives that fast repeatedly, that is a signal to plan smaller traffic allocations or more concurrent tests, not to shorten windows.
go deeper
Be ready to say that reaching the number of users is not the same as covering a representative stretch of calendar, and that three days misses the weekend entirely.
Explain both mechanisms: the day-of-week coverage gap and the fact that heavy users are exposed first, so a fast sample is skewed towards frequent visitors.
Show that the estimate is internally valid but externally limited, and turn the surplus traffic into a plan: smaller allocations, more concurrent tests, or a tighter minimum effect.
Own the portfolio view. Decide how much traffic each test may claim, what the standard runtime is, and what evidence justifies a documented short-window exception.
## Two preconditions, not one A sample-size calculation answers the question "how many units do I need for this design to detect the effect I care about?" It does not answer "which units?" or "over what stretch of calendar?" A runtime decision answers those. A well-formed test plan therefore states both: a target sample and an end date, with the result read when both are satisfied. A high-traffic product routinely hits the count long before the calendar is done, and that is exactly the situation this question probes. ## Problem one: the days you did not cover Three days of a seven-day cycle is a mid-week sample. If the metric or the effect behaves differently at the weekend — different audience, different intent, different device mix — then the estimate describes mid-week users and is being generalised to everyone. The comparison between arms is still internally sound, because both arms drew from the same three days; what is missing is external validity for the population the feature will actually be launched to. ## Problem two: who arrives first This one is easy to miss and is what distinguishes a strong answer. Exposure happens when a user visits. Someone who visits daily is almost certain to be exposed on day one. Someone who visits monthly has roughly a one-in-thirty chance of appearing in a three-day window. So the early sample is dominated by heavy users, and the tail of light, occasional and lapsed users only fills in as the window lengthens. That matters because heavy users are usually not average users: they convert at different rates, have already learned the interface, and often respond differently to a change. A change that helps newcomers find something may look flat among power users who already knew where it was, and vice versa. So even if there were no weekly cycle at all, a very short window would still measure a different population than a two-week window does. ## Problem three: the system may not have settled The opening days of a launch are also when caches are cold, staged rollouts are still widening, and any instrumentation problem is most likely to be live. Reading a result from exactly that stretch is reading the least stable part of the experiment. ## What to actually do - Keep the test running to the committed end date. State in the plan, before launch, that the read happens when both the sample target and the calendar end are met. - If the sample arrives dramatically early and often, that is a capacity signal. Options: allocate a smaller share of traffic to each test so more tests run concurrently, target a smaller minimum effect that the extra sample can now support, or use the surplus to power secondary metrics and pre-declared segment breakdowns. - Do not convert surplus traffic into an ad-hoc shorter runtime after the fact. A duration chosen after seeing the numbers is not a duration you can defend. ## When a short window is genuinely acceptable There are real cases. A metric with no weekly pattern in historical data, a population that is not visit-frequency-heterogeneous, an emergency guardrail check where a fast read on a catastrophic regression is worth more than representativeness, or a within-session interaction metric that is measured immediately at exposure and has no delayed component. In every one of those cases the argument is made from evidence about the metric and stated in the plan in advance — not improvised because the count filled up quickly. ## How to answer it in an interview Say clearly that sample size and runtime are separate requirements, give the day-of-week reason and the heavy-user reason, and then show judgment: if you are consistently over-supplied with traffic, redesign the experiment portfolio around it rather than shortening windows one test at a time.
- You keep hitting the sample target in days. What should change about how you plan experiments?Treat it as spare capacity. Allocate a smaller traffic share per test so several run side by side, or aim at a smaller minimum effect that the extra sample now supports. You can also pre-declare segment cuts and secondary metrics that the surplus makes readable. What you should not do is shorten runtimes test by test.
- Is there a metric for which a three-day window really is enough?Sometimes. An immediate within-session interaction metric, on a product whose historical data shows no weekly pattern and whose users do not vary much in visit frequency, can be read from a short window. The case has to be made from history and written into the plan before launch, not argued after the count fills.
- Does the fast sample at least mean the arms are comparable?Yes, internally. Both arms were drawn from the same three days by the same randomisation, so the contrast is unbiased for those days and that population. The problem is external: that population is skewed towards frequent, mid-week visitors, so the estimate may not carry over to the full rollout audience.
saying these in an interview costs you the question
- Treats the sample count as the only stopping condition
- Says more users always means a more representative sample
- Ignores that early exposure favours the most frequent visitors
- Shortens the window after seeing how fast traffic arrived
- Assumes three days of high volume beats a full week