skip to content

How would you use an interrupted time series to measure a site-wide redesign shipped on 1 March?

level: seniorimportance: should knowfreq 42%

answer

  1. the running variable is the date
  2. extend the pre-trend across the ship date
  3. step change or bend in the trajectory
  4. what else launched that day?
  5. consecutive days are correlated

basics

~20 s

Treat the ship date as a cutoff on calendar time: fit the pre-period trend, then allow both an immediate level shift and a later slope change. The main threat is anything else that shipped that day.

solid answer

~50 s

An interrupted time series is a discontinuity design whose running variable is the date and whose cutoff is 1 March. Fit the daily metric on time, add an indicator for the post period and an interaction with days since 1 March: the indicator estimates the **immediate level shift** and the interaction estimates the **change in slope**. Distinguishing the two matters — a redesign can drop the metric on day one and recover, or leave it flat and change its trajectory, and reporting a single "lift" hides which happened. Restrict the window to a period around the date rather than fitting years of history, exclude a short transition window if the rollout was not instantaneous, and use standard errors robust to correlation between consecutive days, because naive errors on serially correlated data are far too small. The decisive question is not statistical: what else shipped on 1 March?

go deeper

for a junior

Know that the comparison is against the continuation of the pre-launch trend, not against the raw average of the weeks before, and that the launch date plays the role of the cutoff.

for a middle

Be able to write the specification and say what the post indicator and the time interaction each estimate, and why a growing metric makes a simple before-after mean difference misleading.

for a senior

Lead with the confounding audit — release calendar, campaigns, traffic mix — handle a smeared rollout with a transition window, and use standard errors that respect correlation between consecutive days.

for a principal

Own the evidentiary standard for launch claims: decide when a single-series estimate may be called causal, when it must be reported descriptively, and when the product should have held back an untouched slice to compare against.

## The setup When a change is switched on for everyone at a known moment, there is no untreated group to compare against — only the past. An interrupted time series treats calendar time as the running variable and the ship date as the cutoff, so the design logic is the discontinuity logic: fit the trend approaching the date from the left, fit the trend after it, and read what changes at the boundary. A workable specification on daily data, with `t` counting days and `T0` the ship date: `Y_t = a + b*t + c*Post_t + d*(t - T0)*Post_t + error` - `Post_t` is 1 from 1 March onward, 0 before. - `c` is the **level shift**: the immediate jump in the metric on the ship date, holding the pre-existing trend constant. - `d` is the **slope change**: how much the day-over-day trajectory changed after the redesign. - `b` is the pre-period trend, which the design credits to whatever was already happening rather than to the redesign. ## Level shift versus slope change These are genuinely different stories and interviewers want them separated. - **Level shift only** (`c` non-zero, `d` about zero): the metric stepped to a new plateau. A redesign that removes a step from a checkout flow looks like this. - **Slope change only** (`c` about zero, `d` non-zero): nothing happened on day one, but the trajectory bent. A change that improves discoverability as users learn the new layout looks like this. - **Both, with opposite signs**: the classic novelty or disruption pattern — an immediate drop as users relearn the interface, followed by a steeper recovery. Reporting a single averaged "effect" over the post period would show roughly nothing and conclude, wrongly, that the redesign did not matter. Always plot the raw series with the fitted pre-trend extended across the ship date. The gap between the extrapolated pre-trend and the observed post data is the visual version of the estimate, and if it is not visible in the picture, be sceptical. ## What has to be true The design assumes the metric would have continued along its pre-period path had the redesign not shipped. Three things attack that. 1. **Co-occurring changes.** If a pricing change, a marketing push, or an infrastructure migration also landed on 1 March, the level shift is theirs as much as the redesign's, and no amount of modelling separates them. Reconstruct the deployment and campaign calendar around the date before estimating anything. This is the first question to ask and the most common reason the design fails. 2. **Anticipation and gradual rollout.** If the redesign reached users over a week, or an announcement changed behaviour beforehand, the cutoff is smeared. Exclude a transition window on both sides and say how wide it was. 3. **Composition shifts.** If traffic mix changed at the same time — a new acquisition source, a different device split — the metric can move without any user behaving differently. Check that the mix is stable across the date, and look at the metric within stable segments. ## Inference Consecutive days of the same metric are correlated, so residuals are not independent. Naive standard errors on such data are badly understated and will declare small shifts significant. Use standard errors robust to serial correlation, or aggregate to a coarser interval, and be conservative about a marginal result. Also beware short windows: a few weeks on each side of the cutoff give very little information about a trend, and a single anomalous day can drive the whole estimate. ## Why this is weaker than a cross-sectional discontinuity In an eligibility-cutoff design, units just below and just above the threshold are different units that are otherwise alike, and you can argue that nothing else changes at the threshold. With time as the running variable, the units just before and just after are the same population at different moments, and the world genuinely does change over calendar time. There is also no manipulation logic to appeal to — the date is known in advance to everyone inside the company, and behaviour around a launch is anything but incidental. The practical consequence is that the estimate deserves weaker language. Say "the metric stepped up by 3% on the ship date and held" rather than "the redesign caused a 3% lift", unless the co-occurrence check came back clean and the pattern is large relative to the series' normal variation. Where a slice of the product was left untouched, comparing against it is a stronger design and worth proposing. ## What a strong answer contains Name the specification and what each term means, separate level from slope, propose the plot, and lead with the co-occurring-changes audit. Candidates who go straight to a significance test on a post-versus-pre mean difference have skipped both the trend and the question of what else happened that day.

  • Why not simply compare the mean metric in the month before and the month after?
    Because it credits the pre-existing trend to the change. A metric already growing 1% a week will show a post-period gain with no intervention at all. Fitting the trend and reading the departure from it separates growth that was already happening from the step or bend at the ship date.
  • The metric drops on 1 March and then climbs steeply. What do you report?
    Both terms and the picture: an immediate negative level shift with a positive slope change, consistent with users relearning the interface. Report when the series crosses back over the extrapolated pre-trend, and hold the call until enough post-period days exist to distinguish recovery from a permanent loss.
  • How does time as a running variable weaken the design compared with an eligibility cutoff?
    An eligibility cutoff separates different units that are otherwise alike, and you can argue nothing else changes at that score. Time separates the same population at different moments, during which the world moves independently of your change, and the date is known in advance to everyone shipping other things that week.
  • What would make you refuse to give a causal number at all?
    A crowded release calendar around the date, a rollout spread over days with no clean cutoff, a traffic-mix change at the same time, or a series so volatile that the shift is inside its normal range. In those cases report the observed movement descriptively and propose a design with a comparison group.

saying these in an interview costs you the question

  • Compares post-period mean to pre-period mean, ignoring the trend
  • Reports one lift number without separating level from slope
  • Never audits what else shipped on the same date
  • Uses naive standard errors on daily correlated data
  • Calls a novelty dip a permanent regression after three days

context