skip to content

A planned two-week test window overlaps Black Friday. How do you handle the runtime?

level: seniorimportance: should knowfreq 44%

answer

  1. not just more traffic, different traffic
  2. who visits and why both change
  3. baseline and variance both move
  4. decide at planning time, not mid-flight
  5. shift the window or pre-register

basics

~20 s

Decide before launch. The cleanest option is usually to shift the window so it sits entirely before or after the holiday. If the test must span it, keep whole weeks and treat the result as an estimate for holiday traffic.

solid answer

~50 s

The key move is that this is a launch-time decision, not an in-flight one. Options, roughly in order of preference: shift the window to sit wholly before or wholly after the holiday period; or, if the calendar forces it, run through and pre-register how you will read it. A holiday week is not just a traffic spike — the visitor mix changes, intent changes, the baseline conversion rate and its variance change, and competing promotions and marketing pushes land at the same time. That has two consequences. Sizing assumptions built on normal-week baselines no longer hold, so the extra volume does not simply buy extra power. And the effect you measure is the holiday effect, which may not carry to a normal week. If the feature is itself holiday-facing, that may be exactly the estimate you want. Otherwise, plan the window around it and say so in the test plan.

go deeper

for a junior

Be ready to say that a holiday week brings a different kind of visitor, not just more of them, and that the test calendar should be checked against known events before launch.

for a middle

Explain why the baseline rate and variance shift, so extra volume does not automatically mean a better-powered test, and why a pooled holiday-plus-normal window describes neither regime.

for a senior

Show the decision made at planning time: shift, split, run through with a pre-registered separate read, or postpone, and tie the choice to whether the feature is holiday-facing.

for a principal

Own the calendar policy: peak-period experiment freezes, a shared launch calendar across teams, and how much experiment throughput the organisation is willing to trade for clean peak-season operation.

## Why a holiday week is different in kind A holiday or major promotional period is not a scaled-up ordinary week. Four things move at once: 1. **Volume.** Traffic can be several times normal, which fills the sample target quickly. 2. **Audience composition.** Deal-seekers, gift buyers and once-a-year visitors arrive, so the population differs from the one you usually serve. 3. **Behaviour and baseline.** Intent, basket composition, conversion rate and its variance all shift. Baselines used in the sizing calculation stop describing the traffic actually arriving. 4. **Confounding activity.** Marketing pushes, discounts, site-wide banners, inventory constraints and other teams' launches cluster in the same window. The first is why people assume a holiday week is a good time to test. The other three are why it usually is not. ## The options, in order **Shift the window.** If the test can start earlier and finish before the period, or start after it, do that. It is the only option that leaves you with a clean estimate for a normal week, and it costs nothing but scheduling. **Run before and after with a gap.** Occasionally the only clean traffic sits on both sides. Splitting one experiment across a discontinuity is fragile — assignment persistence, changing populations, a different product surface — so prefer a shifted window if you can get one. **Run through it, deliberately.** Sometimes the deadline is immovable, or the change is precisely a holiday feature. Then run through, but decide in advance: keep the runtime a whole number of weeks so the day mix stays balanced, pre-register that the holiday stretch will be reported separately as well as pooled, and state in the plan that the headline estimate describes holiday traffic. **Freeze.** Many organisations simply bar launches and experiments during peak commercial periods, because the risk of a bad change during the highest-revenue week outweighs the value of any single test result. That is a legitimate and common answer. ## Things that go wrong - **Assuming the traffic surge buys power.** Power depends on baseline rate, variance and effect size, not on raw counts alone. Holiday traffic has a different baseline and often a much larger variance in revenue-shaped metrics, so a bigger sample can still leave you no better placed to detect a normal-week effect. - **Extending the window mid-flight to "cover" the anomaly.** Changing the runtime after launch because of what the numbers did is a decision made on the outcome. If the calendar hazard was foreseeable, and a national holiday always is, it should have been handled in the plan. - **Reading the pooled number as a normal-week effect.** Pooling one holiday week with one ordinary week gives an average over two very different regimes and describes neither well. - **Ignoring what else shipped.** A holiday period is when marketing spend, promotions and other launches concentrate. Randomisation still protects the contrast between arms, but interaction between your change and a site-wide promotion is real and worth flagging. ## The judgment an interviewer is listening for Three things. First, that you spotted the hazard before launch rather than after. Second, that you can articulate why a holiday week changes population and baseline, not merely volume. Third, that you tie the decision back to the question being asked: if the change is a holiday feature, holiday traffic is the right sample and you should say so; if it is a permanent change to an everyday flow, you want an ordinary window and should move the test rather than the interpretation. ## A practical checklist Before launch, put the calendar next to the proposed window and mark holidays, promotions, planned marketing pushes and known launches from other teams. If any of them lands inside, decide then: shift, split, run through with a pre-registered separate read, or postpone. Write the choice and its reason into the plan. That single step converts an argument you would otherwise have during analysis into a decision nobody has to relitigate.

  • The holiday week triples traffic. Does that not make the test better powered?
    Not straightforwardly. Power depends on baseline rate, metric variance and the effect size, and all three can move during a holiday. Revenue-shaped metrics in particular get much heavier-tailed. You may end up well powered to detect a holiday effect and no better placed to estimate the normal-week effect the decision actually needs.
  • You only notice the overlap on day four of a running test. What now?
    Do not silently re-plan. Options are to stop and rerun in a clean window if the change is not urgent, or to continue to the committed end date and report the holiday stretch separately alongside the pooled result, flagging the limitation. Whichever you pick, record that the decision was made mid-flight so readers can weigh it.
  • When is running through the holiday actually the right call?
    When the holiday population is the population the decision is about — a gifting flow, a promotional banner, peak-season checkout capacity. Then holiday traffic is the relevant sample rather than a contaminant, and the limitation is reversed: the result should not be extrapolated to ordinary weeks.

saying these in an interview costs you the question

  • Assumes a holiday traffic spike simply improves power
  • Extends or truncates the window mid-flight to dodge the anomaly
  • Reports a pooled holiday and normal week as a typical-week effect
  • Only notices the calendar clash during analysis
  • Ignores concurrent promotions and other teams' launches

context