When is a switchback design, alternating treatment on time slices, better than a user-level test?
answer
- shared supply, not independent users
- the whole market flips together
- the time slice is the unit
- carryover across the boundary
- count slices, not rides
basics
~20 sUse it when interference is system-wide — a shared pool of drivers or inventory that no user-level split can separate. Alternating the whole market between conditions on short slices makes the time slice the randomized unit.
solid answer
~50 sA user-level split fails whenever treatment changes a shared resource. In ride-hailing, a dispatch or pricing change moves the same drivers, so treated and control riders draw from one contaminated pool and neither arm ever sees a clean equilibrium. A switchback puts the entire city into treatment for a slice — say 30 minutes — then into control for the next, randomizing the order, so every slice experiences the full system-level effect and the estimate answers the launch question. The costs are real. Consecutive slices are not independent, because driver positions and open trips carry over, so a burn-in window at each boundary is discarded. The effective sample size is the number of slices, not rides, which means weeks of switching for a modest effect. Slice length trades carryover, which wants long slices, against power, which wants many. Time-of-day effects must be balanced by randomizing within blocks.
go deeper
Know what a switchback is: the whole market is switched between treatment and control over successive time windows instead of splitting users into two simultaneous groups.
Explain why a shared driver or inventory pool invalidates a user-level split, and why the randomized unit becomes the time slice rather than the individual rider or trip.
Demonstrate operating one: block randomization across hours, a burn-in discard at boundaries, slice-level aggregation, and a power calculation based on slice count and weeks of running.
Own when weeks of market-wide switching are worth it versus a fast biased readout, and set the rule for which classes of change may never ship on a user-level test.
## The failure a switchback is built for Some treatments do not act on a user; they act on a **system**. A change to how a ride-hailing platform matches requests to drivers, or how it prices during a surge, changes where drivers are and how long they stay busy. Those drivers then serve everyone. If you split riders into treatment and control, a treated rider's match consumes a driver who is no longer available to a control rider — the arms are coupled through the supply pool, and the control arm is no longer a picture of the world without the change. The symptom is that no amount of care in the user-level design helps. The split can be perfect, the metrics well-instrumented, the sample enormous, and the answer still wrong, because the quantity being compared is not "system with the change" versus "system without it" but two halves of a single hybrid system. ## The design A switchback randomizes **time**. The whole market runs under one condition for a fixed window, then the coin is flipped again for the next window. Within a slice, everyone experiences a coherent world: the dispatch logic is uniform, the supply equilibrium reflects that logic, and the metrics for that slice describe the system under that condition. The treatment effect is estimated by comparing aggregate outcomes of treatment slices against control slices. This is why it works where user splits do not: there is no cross-arm contamination *within* a slice, because there is no other arm running at the same time. ## The four costs **1. Carryover across boundaries.** The market has memory. When the condition flips at the top of a slice, drivers are still positioned where the previous condition put them, and trips started under the old regime are still running. The first part of each slice is a blend. The standard remedy is a **burn-in discard**: drop the first several minutes of each slice from the analysis and measure only the settled portion. That costs data but protects the comparison. **2. Slice length is a genuine trade.** Longer slices give the system time to settle and make the discarded burn-in a smaller share of the window, reducing carryover contamination. Shorter slices give more slices, which is where precision comes from. Neither end is safe: very long slices produce a handful of units and no power; very short slices measure mostly transition states. **3. The sample size is the slice count.** Rides inside a slice share an assignment and the same market state, so they are correlated and are not independent observations. Treating individual rides as the unit produces standard errors that are far too small and confidence intervals that are dramatically overconfident. The defensible analysis aggregates each slice to one number — the slice's average outcome — and compares those across arms, which makes precision a function of how many slices you ran and how much slices vary. Since a day contains only a few dozen usable 30-minute slices, a switchback usually needs weeks. **4. Time is a confounder.** Demand, weather, driver supply and rider mix all vary sharply by hour and by weekday. With a limited number of slices, a plain coin flip can leave treatment overrepresented in rush hour. The fix is **block randomization**: within each short block of adjacent windows, assign one to treatment and one to control, so the arms are balanced across time of day and day of week by construction rather than by luck. ## When it is the wrong tool A switchback is expensive and slow, and it is the wrong choice when: - **Units genuinely do not interact.** If the change is a rendering tweak with no effect on the shared pool, a user-level split gives an unbiased answer with far more power. Do not pay for interference protection you do not need. - **The effect takes days to materialise.** A switchback measures what happens inside a slice. Learning effects, habit formation and anything that unfolds over a longer horizon cannot be read from 30-minute windows. - **Carryover is long relative to any workable slice.** If the market takes an hour to settle, honest slices become hours long, and you lose the slice count that makes the design work. ## Diagnostics worth running Check whether the estimate is stable across slice lengths — a large sensitivity to slice length is a sign that carryover is not contained. Check that the treatment share is balanced within hour-of-day and day-of-week buckets. Compare the estimate with and without the burn-in discard: a big gap says the boundaries are dominating the measurement. ## What a strong answer contains Name the shared resource that makes the user split invalid, describe the alternation and why each slice is internally coherent, then volunteer the costs unprompted — carryover and burn-in, slice count as sample size, block randomization against time confounding, and the weeks of running that all this implies. A candidate who describes the design without pricing it has not operated one.
- How do you choose the slice length?Trade carryover against power. A slice must be long enough for the system to settle and for the discarded burn-in to be a small share of it, but short enough to yield many slices, since precision comes from slice count. Check that the estimate is not sensitive to the length you picked; strong sensitivity means carryover is still contaminating the comparison.
- Why can you not treat every ride in a switchback as an independent observation?Rides inside one slice share the same assignment and the same market state, so they are correlated. Counting them individually understates the variance and produces confidence intervals that are far too narrow, along with a false-positive rate well above the nominal level. Aggregate each slice to one observation and let the number of slices drive precision.
- What confounder does a switchback introduce that a user-level test does not?Time. Demand, weather and driver supply swing by hour and weekday, so with few slices an unlucky draw can align treatment with rush hour. Randomize within blocks — one treatment and one control slice inside each pair of adjacent windows — so the arms are balanced across time of day by construction.
saying these in an interview costs you the question
- Counts individual rides as independent observations
- Ignores carryover between consecutive slices
- Picks a slice length with no burn-in discard
- Assumes randomization handles time of day automatically
- Runs a switchback where units never interact