skip to content

When should an experiment's opening days be discarded as a burn-in period?

level: seniorimportance: nice to knowfreq 30%

answer

  1. the system, not the users, was unsettled
  2. cold caches and an unfinished rollout
  3. declared before launch, not after the read
  4. the criterion must not mention the metric
  5. add the discarded days back to the end

basics

~20 s

Discard opening days only when the system, not the users, was unsettled: cold caches, an incomplete rollout, models still retraining. Declare the burn-in and its criterion before launch, and extend the runtime to replace discarded days.

solid answer

~50 s

A burn-in is a pre-declared stretch at the start of an experiment whose data you exclude because the serving system was not yet in its steady state. Legitimate causes are mechanical: caches warming so the variant was artificially slow, a staged rollout that had not reached full exposure, a model still retraining on freshly logged behaviour, or a logging bug fixed on day two. In all of those the arm you measured is not the arm you intend to ship. The discipline is what makes it defensible. Write the burn-in length and the readiness criterion into the plan before launch, base the criterion on system state rather than on the metric, and add the discarded days back onto the end so the runtime still covers whole weeks. Dropping the first two days after seeing they hurt your result is not a burn-in; it is an outcome-dependent exclusion.

go deeper

for a junior

Be ready to say that the first days of a launch can reflect cold caches or an unfinished rollout rather than real user response, and that any exclusion has to be planned in advance.

for a middle

Explain the mechanical causes and the rule that the burn-in criterion is about system readiness, never about how the metric is trending. Note that both arms lose the same window.

for a senior

Demonstrate the discipline: pre-registered length or criterion, symmetric exclusion, runtime extended to replace discarded days, and results reported with and without the burn-in.

for a principal

Own the prevention side. Push for a warm-up step before exposure begins and a standing rule on when exclusions are permitted, so individual teams are not negotiating it per test.

## What a burn-in is Burn-in is the deliberate exclusion of an experiment's opening window from analysis, on the grounds that the system under test had not reached steady state. It is a statement about infrastructure readiness, not about user behaviour. ## Legitimate reasons - **Cold caches and cold paths.** A new code path serves from an empty cache at first. Latency in the variant can be materially worse on day one for reasons that will never exist in the steady state, and latency drags conversion metrics with it. - **Incomplete rollout.** Deployments propagate across regions, instances or app versions. Until propagation completes, the variant population is not the intended population. - **Components that learn.** A ranking or recommendation stage that trains on logged interactions needs a stretch of the variant's own traffic before its output is representative of what it will do in production. - **Instrumentation being repaired.** A logging or assignment bug found and fixed on day two makes the pre-fix data untrustworthy for both arms. In each case the honest description is: during that window, the thing being measured was not the thing you plan to ship. ## What makes it defensible The difference between a burn-in and a convenient exclusion is entirely in the ordering and the criterion. 1. **Pre-declare it.** The plan names a burn-in length, or names an explicit readiness criterion, before any user is exposed. "We exclude the first 48 hours" is a plan. "We excluded the first 48 hours" written after the read is not. 2. **Make the criterion mechanical.** Cache hit rate above its normal band; rollout at one hundred percent of targeted hosts; the retraining job having completed one full cycle; assignment logs passing their integrity check. None of these mention the experiment's outcome metric. A criterion phrased in terms of the metric — "until the difference stabilises" — makes the exclusion a function of the result and destroys the guarantee. 3. **Symmetry.** The same window is dropped from both arms. Dropping only the variant's bad opening days is not an analysis choice, it is fabrication. 4. **Extend, do not shrink.** If two days are discarded from a 14-day plan, run 16 so the analysed window still covers whole weeks and still meets the sample target. Discarding days without extending quietly turns a balanced two-week design into an unbalanced 12-day one. 5. **Report both.** Show the result with and without the burn-in. If they agree, the exclusion was harmless and the reader is reassured. If they disagree sharply, that difference is itself a finding worth explaining. ## The cleanest alternative Burn-in inside an experiment is a second-best fix. The better move, when you can afford it, is to warm the system before the experiment starts: deploy the code, let caches fill and the rollout complete under a small holdout or a no-op configuration, and only then begin exposure and start the clock. Then there is nothing to discard, no exclusion to justify, and the whole window is analysable. ## How to talk about it An interviewer is usually testing whether you can distinguish a principled, pre-registered exclusion from post-hoc data pruning. Lead with the mechanical causes, state that the criterion must be about system readiness rather than about the metric, mention symmetry across arms, and say that discarded days are added back to the end of the runtime rather than subtracted from it. Mentioning that you would report the analysis both ways is what a careful practitioner adds and what a rushed one leaves out.

  • What is wrong with deciding the burn-in length by watching when the treatment effect stabilises?
    It makes the exclusion a function of the outcome. Any wobble in the opening days can be relabelled as not-yet-settled until the remaining window shows the result you expected, so the analysis stops being a test of anything. A burn-in criterion has to be about system readiness — rollout completion, cache state, a finished retraining cycle — and has to be written down first.
  • If two days are discarded from a planned 14-day test, what happens to the end date?
    Push it out by two days. Otherwise the analysed window is 12 days, which is no longer a whole multiple of seven and no longer covers each weekday equally, and it falls short of the sample target the plan was built on. Extending keeps both the calendar balance and the power intact.
  • How can you avoid needing a burn-in at all?
    Warm the system before exposure starts: deploy the code, let the rollout complete and caches fill under a no-op or holdout configuration, verify the assignment and logging pipeline, and only then start exposing users and counting days. Nothing has to be excluded, and there is no exclusion decision to defend later.

saying these in an interview costs you the question

  • Drops the first days only after seeing they hurt the result
  • Defines the burn-in by when the metric looks stable
  • Excludes the opening window from one arm only
  • Discards days without extending the end date
  • Treats burn-in as a routine default with no stated cause

context