skip to content

Progressive Rollouts & Canary Analysis

Exposing changes to a small slice of traffic and judging them against SLI baselines before going wide. Interviewers ask how you would catch a bad deploy before users do — canary analysis and staged rollouts are the answer.

on this pageshow

questions

6

You roll a change out to 1% of production traffic, the canary's error rate and latency stay clean for an hour, and then the change breaks one enterprise customer's integration the moment it goes to 100%. Why can a 1% traffic canary contain 0% of the users a change affects, and how would you choose the canary population instead?

level: middleimportance: must knowfreq 60%

answer

  1. think about who, not how many
  2. rare paths contribute very few requests
  3. one tenant may send zero canary requests
  4. 0.99^20 is about 82 percent
  5. select the cohort, then sample randomly

basics

~20 s

A traffic percentage is a random sample, not a representative one. A rare code path, a single tenant, or one old client version can send so few requests that the canary receives none of them. Choose the canary population deliberately, not just its size.

solid answer

~60 s

Routing 1% of requests picks requests at random, and randomness is only representative when the thing you are testing is spread evenly across requests. If a customer sends twenty requests in the canary window, the chance none of them lands on the canary is `0.99^20`, about 82% — the canary is genuinely blind to that customer. The same holds for a code path that is 0.5% of traffic, a locale, or a mobile client version that only a slow-upgrading minority still runs. So I size the canary by how many events of the *risky path* it needs to see, not by a round percentage, and where the risk is concentrated I select the population instead of sampling it: internal and dogfood traffic first, then a named cohort of tolerant tenants, then random traffic, and the highest-volume or highest-value customers last. For anything stateful I hash on user or tenant ID so a session does not flip versions mid-flow, and I make the assignment a label on the telemetry so I can actually attribute an error to the canary.

code

python · 10 lines
python
canary_share = 0.01
requests_from_customer = 20
p_never_hits_canary = (1 - canary_share) ** requests_from_customer
print(round(p_never_hits_canary, 3))

rps = 1000
path_share = 0.005          # the changed code path's share of traffic
bake_seconds = 15 * 60
samples = rps * canary_share * path_share * bake_seconds
print(round(samples))

go deeper

for a junior

Know that a canary sends a small share of real production traffic to the new version before everyone gets it, and be able to say why that is safer than switching everything at once.

for a middle

Be ready to explain that a traffic percentage is a random sample and to do the arithmetic out loud: how many requests a 1% canary gives the specific code path you changed, and how likely it is to miss a low-volume customer entirely.

for a senior

Show that you design the population, not just the percentage — internal traffic, then a chosen cohort, then random share, with the highest-volume customers staged last — and that you label telemetry by version so canary errors are attributable at all.

for a principal

Own the tradeoff between early signal and representative signal: deliberately chosen cohorts find concentrated risks fast but are unrepresentative, while random shares are representative but blind to the tail. Be able to say what your release policy requires for each risk class and who maintains the cohorts.

## The claim a percentage rollout is really making "1% of traffic is on the new version" is a statement about request routing. The statement people *hear* is "we have tested the change against 1% of reality". Those are only the same claim when the failure you are looking for is spread uniformly across requests — a change that breaks every request, or slows down the hottest endpoint. For anything narrower, a percentage of traffic is a sample, and a small random sample of a skewed population misses the tail almost every time. ## The arithmetic, because it is the whole point Two numbers decide whether a canary can see a thing. **Coverage of a cohort.** If a customer sends `n` requests during the canary window and each request has probability `p` of being routed to the canary, the chance the canary never sees that customer is `(1 - p)^n`. At `p = 1%` and `n = 20`, that is `0.99^20 ≈ 0.82`. Four times out of five you learn nothing about them. To be reasonably confident of at least one hit you need `n` in the hundreds at 1%. **Sample count for a rare path.** Take a service at 1,000 requests per second. A 1% canary receives 10 rps. If the changed code path is 0.5% of traffic, the canary exercises it 0.05 times per second — three times a minute. A fifteen-minute bake gives about 45 executions, which cannot distinguish a jump from a 0.1% error rate to a 1% error rate from noise. The canary is green because it is empty, not because it is healthy. ## Choosing the population instead of sampling it The fix is to treat "who is in the canary" as a design decision with its own ladder, roughly: 1. **Synthetic and internal traffic.** Cheap, fully attributable, and you own the complaint. It catches gross breakage before any customer is exposed. 2. **Dogfood / employee accounts.** Real workflows, real data shapes, a population that tolerates breakage and reports it in words rather than in a support ticket. 3. **A named cohort.** Tenants or accounts chosen because they exercise the changed path, or because their tolerance is known. This is what you use when the risk is concentrated — the integration partner whose contract you touched should be *in* the canary, deliberately, not sampled into it with probability 0.18. 4. **Random traffic share.** Now the percentage does its real job: measuring aggregate SLIs on a representative slice. 5. **The whale last.** If one customer is 70% of your volume, any canary containing them is not a canary, and any canary excluding them tells you nothing about them. Stage them as their own step, with their own bake and their own watchers. ## Per-request versus per-entity assignment Random per-request routing is fine for a stateless read. It is actively harmful for a multi-step flow: a checkout whose first call hits the new version and whose second call hits the old one exercises a version combination you never intended to test, and the resulting error is attributable to neither. Hash on a stable key — user ID, tenant ID, session — so an entity stays on one version for the whole rollout. That is also what makes the analysis honest: you can count *users* affected, not just requests. ```text canary_bucket = hash(tenant_id) % 1000 < 10 # sticky 1% by tenant, not by request ``` The cost of sticky assignment is that your 1% of *entities* is not 1% of *traffic* — if the hash happens to catch a heavy tenant, the canary may carry 8% of requests. Measure the realised share rather than assuming the configured one. ## Attribution None of this works if you cannot tell canary telemetry from baseline telemetry. Every request the canary serves should carry the version and the assignment as a label so error rate, latency and business metrics can be sliced by it. Without that label a 1% canary contributes 1% of the errors to an aggregate dashboard, where a total failure of the new version looks like a 1% blip and is invisible against normal noise. ## What this costs Deliberate populations are more work than a percentage: someone maintains the cohort, cohorts go stale, and a cohort of tolerant customers is by definition unrepresentative of the intolerant ones. The honest position is that population selection buys you *early* signal about specific risks, and the random share buys you *representative* signal about aggregate ones. A serious rollout uses both, in that order.

  • If you assign the canary by hashing the tenant ID rather than per request, what new problem do you have to watch for?
    The realised traffic share stops matching the configured entity share. Hashing 1% of tenants into the canary can send it 8% of requests if a heavy tenant lands in the bucket, or almost none if only idle tenants do. Measure the actual request share the canary receives and size your capacity and your analysis from that, not from the configured percentage.
  • One customer is 70% of your traffic. How do you canary at all?
    Accept that they are their own rollout stage. Everything else — internal, dogfood, the long tail of small tenants — goes first and validates the change generally. Then the whale gets a dedicated stage with its own bake, its own watchers and its own abort criteria, ideally starting with a fraction of their traffic if their client supports it. Treating them as 70% of a random sample means every canary is either blind or full-blast.
  • How do you keep a canary cohort from going stale?
    Define it by a property rather than by a list of account IDs — tenants on a given plan, accounts that called the endpoint you are changing in the last week, or employees. A hand-maintained list drifts until it contains customers who churned and misses the integration that was built last quarter. Re-derive the cohort at rollout time and log which entities it resolved to.

saying these in an interview costs you the question

  • Thinks 1% of traffic means 1% of users are affected
  • Assumes a green canary proves the change is safe
  • Sizes the canary by cost or by a round number only
  • Routes per request even for multi-step stateful flows
  • Cannot slice telemetry by version, so canary errors vanish in the aggregate

context

open as a page

A change passed every canary stage, reached 100% of the fleet, and the failure surfaced two days later. Which classes of failure does a progressive rollout structurally fail to catch, and what controls would you add for them?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Progressive rollouts miss anything that needs time, scale, a specific event, or a specific population: slow leaks and disk fill, dependencies that only saturate at full traffic, monthly batches, and one tenant or old client. Each needs its own control.

open as a page

An automated canary check scores the new version by comparing its error rate and latency against the metrics the rest of the production fleet reported over the same hour. Why is that comparison misleading, and what should the canary be compared against instead?

level: middleimportance: should knowfreq 45%

basics

~20 s

The fleet has been running for days; the canary started minutes ago. Cold caches, unfilled connection pools and runtime warm-up make a healthy canary look slow. Compare it against a control running the old version, started at the same time on the same class of host.

open as a page

You are setting the bake time for each stage of a canary rollout — how long the new version soaks at 1% before it advances to 10%. How do you decide that number, and why is a 10-minute bake worthless for some changes?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Two clocks set the bake: how long it takes to collect enough events at that traffic share to detect the regression you care about, and how long the failure mode you fear needs to appear. Bake for the longer of the two.

open as a page

You set release policy for a service that runs in three regions with three availability zones each. After the initial canary passes, how would you order the rollout across those zones and regions, and how do you decide the total wall-clock time a rollout is allowed to take?

level: principalimportance: should knowfreq 40%

basics

~20 s

Expand along a blast-radius ladder: one zone in the least critical region, then that region, then the remaining regions one at a time, keeping a known-good region until last. Total duration is bounded by mixed-version tolerance, emergency-fix latency and how often you deploy.

open as a page

Your automated canary analysis aborts roughly one rollout in three, nobody can reproduce a defect in the aborted changes afterwards, and engineers are now asking for a flag to skip the check. What is going wrong, and how do you fix it without going blind?

level: seniorimportance: nice to knowfreq 35%

basics

~20 s

The gate is scoring noise: too many metrics, too few samples, and thresholds tuned tighter than normal variance. Fix precision — require a minimum sample count, score a small set of high-signal SLIs relative to a control, and return inconclusive instead of aborting.

open as a page