skip to content

Ramp Plans & Rollout Decisions

Ramping a risky change from 1% to 100%: what each stage is for, why a 1% readout is underpowered and exaggerates surviving effects, and whether phases may be pooled. Ship calls hang on it.

on this pageshow

questions

5

Why ramp an A/B test through 1%, 5%, 25% and 50% instead of starting at full traffic?

level: juniorimportance: must knowfreq 62%

answer

  1. blast radius before statistics
  2. each stage licenses a different decision
  3. the first stage is a smoke test
  4. guardrails and rollback rule per stage
  5. precision arrives only at the large stages

basics

~20 s

A staged ramp limits blast radius: the first small stage exposes few users, so crashes, latency spikes or a collapsing metric are caught cheaply. Later, larger stages exist to measure the effect once the change has been shown to be safe.

solid answer

~50 s

Each stage of a ramp answers a different question. The 1% stage is a safety gate: does the feature break anything — errors, crashes, latency, downstream load — and does any guardrail metric move catastrophically? It is a smoke test, not a read on the business metric, because at that traffic the estimate is far too noisy to settle a small difference. The 25% and 50% stages are where the comparison becomes precise enough to support a launch decision. Ramping also spreads operational risk: infrastructure, caches and downstream services meet the new load gradually rather than all at once. The discipline that makes it work is agreeing, before the ramp starts, what each stage must show to be promoted and what triggers an immediate rollback — otherwise a ramp is just a slow launch with extra meetings.

go deeper

for a junior

Be ready to say what a ramp is and name the first stage's job: catch breakage cheaply. Know that the earliest stage is the noisiest and should not decide a launch.

for a middle

Explain what each stage licenses and why precision only arrives at the larger fractions. Expect to be asked which guardrails you watch and how rollback is actually triggered.

for a senior

An interviewer expects you to design the plan: fractions, promotion criteria, guardrail thresholds, who owns each gate, and how the control arm is constructed so the comparison stays clean.

for a principal

Own the tradeoff between ramp speed and exposure risk across a whole release process, and be able to argue when a ramp is theatre — stages that license no decision are calendar cost you should cut.

## What a ramp plan is A ramp plan (or rollout schedule) is a pre-agreed sequence of traffic fractions at which a change is exposed to randomized users: for example 1%, then 5%, then 25%, then 50%, with the remaining traffic staying on the existing experience as control. It is a *plan*, written before launch, that names the fraction at each stage, what has to be true to move to the next one, and what causes a rollback. ## Why not go straight to full traffic Three distinct reasons, and it helps to keep them separate because they are satisfied by different stages. **1. Blast radius.** If the change is badly broken — a null-pointer on a rare code path, a payment flow that silently fails, a page that never loads on one browser — the number of users harmed is proportional to the fraction exposed. At 1% of traffic, a serious defect is visible in error logs and crash rates within minutes while affecting a hundredth of the population. This is the dominant reason for the first stage, and it is an engineering argument, not a statistical one. **2. Operational risk.** New code paths change infrastructure behaviour: extra queries per request, cold caches, a new downstream call, larger payloads. Ramping lets capacity and latency be observed as load grows, so saturation appears as a warning at 25% rather than an outage at 100%. **3. Measurement.** Only the larger stages give the comparison enough precision to distinguish a modest real improvement from noise. A small early stage can rule out a catastrophe — a metric that halves — but it cannot confirm a two-percent improvement. Treating an early read as a verdict is the single most common ramp mistake. ## What each stage licenses The useful mental model is that a stage *licenses a decision*, and different stages license different ones. - **1%** licenses: continue, or roll back. It answers "is this obviously broken?" Guardrails are error rate, crash-free sessions, latency percentiles, and a sanity check that both arms are logging at all and are roughly the sizes you asked for. - **5%** licenses: continue, or roll back on a guardrail regression that was too small to see at 1%. It also gives the first look at rarer flows — checkout, refunds, an unusual locale. - **25% / 50%** license: the actual launch decision on the primary metric, together with a usable interval around the estimated effect. A stage that licenses nothing is wasted calendar time. If you cannot say what promoting from this stage means, collapse it into its neighbour. ## Guardrails and rollback rules A ramp is only a safety mechanism if somebody is watching and rollback is cheap. Practically, that means: the change sits behind a flag that can be flipped off without a deploy; guardrail thresholds are written down before the ramp ("roll back if p99 latency rises more than 15%, or crash-free sessions fall below the release baseline"); and one named person owns the promotion decision at each stage. Without those, a ramp gives you the illusion of caution. ## Where a ramp does not help A ramp protects against harms that are fast and visible at small exposure. It does not protect against harms that are slow, harms that only appear at scale (a cache that only thrashes above a load threshold, a queue that only backs up at full traffic), or harms concentrated in a population your early stages never reached. It also does not, by itself, make an early estimate trustworthy: exposing fewer users means a noisier comparison, so the earliest numbers are the least reliable ones you will ever see for this change. ## The control side Ramping the treatment is only half of it. The comparison at each stage needs a control group produced by the same eligibility and triggering logic — users who *would* have received the feature but were randomized away from it. Comparing the exposed 1% against "everyone else" mixes in users who were never eligible, were routed elsewhere, or never reached the surface the feature touches, and any difference then reflects who those users are as much as what the feature does.

  • What do you actually check before promoting from 1% to 5%?
    Engineering guardrails first: error and crash rates, latency percentiles, downstream load, and that the feature's own instrumentation fires in both arms. Then a sanity check that the arms are roughly the sizes you requested and that no primary or guardrail metric has moved catastrophically. You are not looking for a positive result at this stage — you are looking for evidence of breakage, and its absence is all the promotion requires.
  • Why hold an explicitly randomized control at 1% rather than comparing against all remaining traffic?
    The remaining traffic contains users who were never eligible, never reached the surface, or were filtered by different logic. A mirror control drawn through the same eligibility and trigger path guarantees both arms are composed and filtered identically, so a difference reflects the feature. Comparing against everyone else mixes a treatment effect with a composition difference, and you cannot separate them afterwards.
  • What kinds of harm does a staged ramp fail to catch?
    Slow-building harms that need weeks of exposure, effects that only appear at scale such as cache thrash or queue saturation above a load threshold, and harms concentrated in a segment the early stages never reached. A ramp is a filter for fast, visible breakage; it is not evidence that the change is good, and it is not a substitute for measuring the effect at a stage with real traffic.

It is a dress rehearsal schedule: a run-through for the crew, then a preview night, then opening. Early nights exist to catch disasters, not to measure the reviews.

saying these in an interview costs you the question

  • Treats the 1% read as the launch decision
  • Says ramping exists mainly to reduce statistical noise
  • Has no pre-agreed rollback rule at any stage
  • Assumes safe at 1% implies safe at 100%
  • Ramps the treatment with no comparable control arm

context

open as a page

An A/B variant read +12% lift at a 1% ramp and +2% at 50% — what explains the drop?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Most likely the winner's curse. The 1% estimate was extremely noisy, and the feature was promoted because that noisy number looked big. Selecting on a large observed effect inflates it, so the precise later read of +2% is the better estimate, not a decay.

open as a page

Should data from earlier ramp stages be pooled into the final effect estimate of an experiment?

level: seniorimportance: should knowfreq 41%

basics

~10 s

Only when every stage estimated the same thing: unchanged treatment, unchanged eligible population, unchanged metric. Even then, pool by combining per-stage effect estimates, never by summing raw counts across stages with different traffic splits.

open as a page

After shipping a feature, is a permanent 1% holdback kept for a quarter worth its cost?

level: principalimportance: should knowfreq 30%

basics

~20 s

Sometimes, and rarely per feature. A holdback buys continued measurement after the experiment ends, but 1% is insensitive and every holdback forks the product. A single company-level holdback covering many launches usually beats one per feature.

open as a page

When ramping an experiment from 1% to 25%, should the original 1% users keep their assignment?

level: middleimportance: nice to knowfreq 33%

basics

~20 s

Yes, by default. Widen who is eligible while leaving already-assigned users exactly where they are. Re-drawing assignment mid-ramp flips some users between arms, and their later behaviour still carries what they already saw, which contaminates both arms.

open as a page