skip to content

Why do A/B tests usually run for whole weeks rather than an arbitrary number of days?

level: juniorimportance: must knowfreq 72%

answer

  1. traffic is not the same every day
  2. weekday visitors versus weekend visitors
  3. a 10-day window counts three days twice
  4. each weekday should carry equal weight
  5. multiples of seven, fixed before launch

basics

~20 s

User behaviour differs by day of week, so a partial week samples an unrepresentative mix. Whole multiples of seven days weight every weekday equally, so the measured lift describes a typical week, not whichever days happened to be included.

solid answer

~50 s

Almost every consumer metric has a weekly cycle: weekday and weekend visitors differ in who they are, what they do and how often they convert. If a test runs Tuesday to Thursday it never sees a weekend shopper, so whatever lift it measures describes mid-week users only. Running a whole multiple of seven days gives each weekday equal representation. A 10-day test launched on a Monday is worse than a 7-day one in this respect, because Monday, Tuesday and Wednesday appear twice and get double weight. The subtle point: because both arms are randomised concurrently, day-of-week effects hit them equally, so the comparison is not biased by the day mix. What a partial week costs you is generalisation — if the effect itself differs by day, you have estimated it only for the days you sampled — plus extra noise. Fix the runtime in the test plan before launch.

go deeper

for a junior

Be ready to say that behaviour differs by day of week and that whole-week windows give every day equal weight. Knowing that 10 days double-counts three weekdays already puts you ahead.

for a middle

Explain the mechanics: which day mix a window covers, why a multiple of seven balances it, and why alignment to Monday is irrelevant. Expect to be pushed on what a partial week actually damages.

for a senior

Show the balance-versus-generalisation distinction and correct the common error about arms seeing different days. Talk about heterogeneous effects by day and about the runtime being committed in the plan before launch.

for a principal

Own the policy angle: a default runtime standard for the organisation, how it interacts with experiment throughput, and when a documented exception is worth the loss of representativeness.

## The weekly cycle Traffic to almost any consumer product rises and falls on a seven-day cycle. The volume differs by day, and so does the *composition* of the audience: weekday daytime visitors are often at work, browsing on desktop, in short sessions; weekend visitors skew towards longer, leisure-driven sessions and different device mixes. Conversion rate, average order value, session length and retention all typically move with the day of week. A business-to-business product usually shows the mirror image, with weekends nearly empty. That cycle is the reason experiment runtimes are quoted in weeks rather than days. ## What a partial week actually costs It is worth separating two claims that are often mashed together, because interviewers probe exactly here. **It does not bias the treatment-versus-control contrast.** Units are randomised into arms concurrently, so Tuesday's visitors are split between arms by the same mechanism as Saturday's. Both arms therefore see the same day mix, whatever that mix is. Someone who says "the control got the weekend and the variant did not" has misunderstood how simultaneous randomisation works; that failure belongs to designs that compare this week against last week, not to a concurrent A/B test. **It does damage generalisation and precision.** Two real costs remain: 1. *Heterogeneous effects.* If the effect differs by day — a delivery-promise banner that matters far more when people shop for the weekend, a checkout change that matters more on mobile-heavy evenings — then a Tuesday-to-Thursday window estimates the mid-week effect and calls it the effect. Ship on that basis and the realised impact over a full week can be materially different, in either direction. 2. *Population mismatch.* A short window over-represents your most frequent visitors, because heavy users show up first and light users trickle in. That is a different audience from the one the rollout will serve. A third, smaller consideration: day-to-day variance is part of the noise in the estimate, and a lopsided window can make the observed variance unrepresentative of the variance the sizing calculation assumed. ## Whole multiples, not round numbers "Whole week" means a multiple of seven consecutive days, not a window aligned to Monday. A test launched Thursday and stopped after the following Wednesday has covered each weekday exactly once and is perfectly balanced. What breaks balance is a duration that is not a multiple of seven: a 10-day test starting Monday covers Monday, Tuesday and Wednesday twice each and the other four days once, so early-week behaviour carries roughly double weight in the pooled estimate. Ten days sounds more conservative than seven and is in fact less balanced. Fourteen days is the common default because it covers two full cycles, smooths week-to-week wobble, and is long enough to be robust to one odd day. ## Committing before launch The runtime belongs in the test plan alongside the sample target, written down before the first user is exposed. Two things follow from that. First, the plan should state the calendar end date, not just a sample count, so that the sample and the calendar are both preconditions for reading the result. Second, the plan should note any known calendar hazard in the window — a promotion, a holiday, a marketing push — because those decisions are much cheaper before launch than after. ## When the rule bends The rule is a heuristic about metric behaviour, not a law. If you can show from historical data that the metric has no meaningful weekly pattern — some internal tools, some low-frequency flows — a non-multiple window is defensible, though you should still commit to it in advance. Conversely, if the metric has a monthly cycle, such as payday-driven spend, seven days is not enough and the runtime argument scales up accordingly: you want the window to cover whole cycles of whatever periodicity your metric actually has. ## What a strong answer sounds like Name the weekly cycle, explain that whole multiples give each weekday equal weight, and then show the extra layer: concurrent randomisation already balances the day mix across arms, so the real prize is an effect estimate that generalises to a normal week, plus a stable variance. Finish by saying the duration was fixed in the plan before launch.

  • If both arms are randomised concurrently, day-of-week effects hit control and treatment equally, so why does the day mix matter at all?
    Because balance across arms and representativeness of the population are different properties. Concurrent randomisation gives you the first: the contrast is unbiased for the days you sampled. It does not give you the second. If the treatment effect is larger at weekends, a mid-week window estimates a mid-week effect, and the number you extrapolate to a full-week rollout is the wrong number.
  • A test launches on a Thursday. Does it have to run until the following Monday to get a whole week?
    No. Any seven consecutive days cover every weekday exactly once, so Thursday through the following Wednesday is balanced. Alignment to a calendar week boundary is not what matters; a duration that is a multiple of seven days is. Waiting for Monday only delays the launch and adds days that unbalance the window.
  • What if the metric genuinely has no weekly pattern?
    Then the whole-week rule buys much less, and a shorter or non-multiple window can be defended — but back it with historical data showing the metric is flat across days, not with an assumption. Even then, commit to a specific duration in the plan before launch, so the end date is a decision rather than a reaction to the numbers.

Tasting three spoonfuls from the top of a pot tells you about the top of the pot. Stirring through the whole week is what makes the taste representative.

saying these in an interview costs you the question

  • Claims the control arm gets weekend traffic that the variant misses
  • Picks 10 days because it sounds safer than seven
  • Assumes every metric is flat across days without checking history
  • Treats runtime as only a sample count with no calendar
  • Thinks the window must start on a Monday to be balanced

context