skip to content

How do canary rollouts and A/B tests differ in what they control for?

level: principalimportance: should knowfreq 33%

answer

  1. bounding damage versus attributing effect
  2. one percent has no power
  3. hold for a load cycle, not a fraction
  4. randomization kills confounding
  5. safe is not the same as better

basics

~20 s

A canary controls risk: a tiny exposed slice bounds blast radius while you watch for breakage, and it is read as a safety check. An A/B test controls inference: randomized arms and sufficient sample size let you attribute a measured effect to the change.

solid answer

~60 s

They answer different questions and are routinely confused because both send traffic to a new variant. A **canary** is risk control. Hold 1% of traffic on the candidate for a period that covers a real load pattern — 48 hours spanning a weekend peak, say — and watch for the failures that are obvious at small volume: errors, latency, cost per request, malformed outputs, safety breaches. Then ramp 25%, then 100%. One percent of traffic almost never has the power to resolve a fractional business-metric change, and it is not meant to. An **A/B test** is statistical inference. Randomize users into arms, pre-declare a primary metric and the sample size needed, run to that size, and read the effect with a confidence interval. Splits are usually large — often 50/50 — precisely because power comes from the smaller arm. So the sequence is canary first to establish the change is safe, then a powered split to establish it is better. Ramping through canary stages and declaring victory because nothing broke is the common mistake.

go deeper

for a junior

Know that a canary exposes a very small share of users first so problems affect few people, while an A/B test splits users into groups to compare which version performs better.

for a middle

Explain why a 1% slice cannot resolve a small business-metric change — power comes from the smaller arm — and which fast, obvious signals a canary can genuinely act on.

for a senior

Show how the stages compose in a real release path and why canary slices routed by region or client version make poor comparison groups even when volume would be sufficient.

for a principal

Own the policy: which classes of change require causal inference versus only safety evidence, what to do when user-level randomization is impossible, and how much calendar time and exposed traffic the organization should spend on measurement.

## Two mechanisms, two questions Both a canary and an A/B test route some live traffic to a new variant, which is why teams blur them. The questions they answer are different in kind, and a change usually needs both. **Canary rollout asks: will this hurt anyone?** It is a *risk-control* mechanism borrowed from operational deployment practice. Expose a small slice, watch, ramp if clean, roll back instantly if not. Its virtue is that the cost of being wrong is bounded by the exposed fraction and the time you sat there. **A/B testing asks: does this help, and by how much?** It is a *causal-inference* mechanism. Randomization makes the arms comparable so that the difference between them can be attributed to the change; sample size determines the smallest effect you can resolve; a pre-declared primary metric prevents you from finding a winner after the fact. One bounds damage. The other produces a number you can defend. ## Why 1% cannot answer an A/B question The smallest effect an experiment can detect scales with the noise in the metric and inversely with the square root of the sample size in the *smaller* arm. A grocery site's conversion rate is noisy at the daily level; resolving a half-percent relative change takes a lot of sessions. A 1% arm accumulates those sessions roughly fifty times slower than a 50/50 split does, which turns a one-week read into an impractical one. So the metrics a canary can genuinely act on are the ones with large, fast, obvious signals: error rate going from near-zero to non-zero, p95 latency doubling, cost per request tripling, a safety classifier firing at all, output-format validity collapsing. Those show up in hundreds of requests. Conversion does not. The corresponding mistake is the mirror image: running a 50/50 split on a change that has never been exposed at all. If it is broken, you broke it for half your users, and you find out only when the metrics come in. ## What a canary actually controls Beyond sample-size logic, a canary controls things a randomized split does not. **Blast radius over time.** The exposed population is small *and* the exposure is reversible in minutes. That combination is what makes it safe to try a structural change on production at all. **Coverage of load patterns.** A canary is held for a duration, not just a fraction, precisely because failures are often time-dependent — a weekend peak, a nightly batch, a Monday-morning burst, a cache that only fills after hours. Holding 1% for 48 hours across a weekend peak before ramping is a deliberate choice to see the system under a load shape a Tuesday-afternoon test never produces. **Staged confidence.** Each ramp step is a new observation at higher volume. Failure modes with low per-request probability become visible at 25% that were invisible at 1%, so the ramp is itself a sequence of tests rather than a formality. ## What an A/B test controls **Confounding.** Seasonality, marketing pushes, an outage, another team's launch — all hit both arms simultaneously, so a difference between arms is attributable to the change. A before/after comparison, which a canary superficially resembles, has no such protection: the world changed between the two periods too. **Selection.** Random assignment ensures the arms have comparable users. Canary slices are often *not* random — routed by region, by pod, by client version, by internal-user flag — which makes them poor comparison groups even when volume would suffice. **Multiple comparisons and peeking.** A disciplined split pre-registers the metric and the stopping rule. Repeatedly checking an experiment and stopping when it looks good inflates false positives, which is why sequential testing methods exist for teams that need to watch continuously. ## How they compose The usual release path for a meaningful LLM change: offline gate, then shadow on mirrored traffic where side effects allow, then a small canary held long enough to cover a real load cycle, then a powered randomized split on the primary metric with guardrails, then a staged ramp to full exposure. Note that canary and A/B can overlap physically — a 1% canary arm *is* randomized in many implementations, and its data can contribute to the eventual analysis. The distinction is not the plumbing, it is what you are entitled to conclude. Nothing breaking at 1% licenses more exposure; it does not license the claim that the change is better. ## Judgment calls a lead owns **When to skip the split entirely.** Some changes are not worth a powered experiment: a latency optimization with no expected quality change, a security fix, a cost reduction at measured parity. Canary plus guardrails is sufficient, and demanding a full A/B on everything is its own organizational cost. **When a split is impossible.** Small B2B populations, strong network effects between users, and changes that alter shared state often cannot be cleanly randomized at the user level. Then the toolkit shifts — switchback designs, cluster randomization by account or region, interrupted time series — each with weaker inference that you should name honestly rather than pretend a user-level split existed. **How long to hold.** Long enough to cover the load and behaviour cycles that matter, and long enough for novelty and disruption effects to decay. Both arguments push toward days, not hours, for user-facing changes. **How much of the release budget experimentation deserves.** Every powered experiment costs calendar time and holds a fraction of users on the worse variant. A team that A/B-tests every prompt tweak ships slowly; a team that canaries everything and measures nothing accumulates unverified changes. Setting that policy — which classes of change require inference and which only require safety — is the principal-level call. ## What interviewers listen for The expected answer separates risk control from inference and explains why sample size makes a canary structurally unable to settle a fractional business-metric question. Strong candidates add the non-random-slice problem, the role of hold duration, and the cases where randomization is not available at all.

  • When is a canary plus guardrails enough, with no powered split at all?
    When the change is not expected to move the primary metric and the risk is operational: a latency or cost optimization at measured quality parity, a security fix, an infrastructure migration. Demanding a powered experiment there spends weeks to confirm a null. The condition is that you can state in advance what would count as harm and watch it, which the guardrails do.
  • What do you do when user-level randomization is not possible?
    Shift design rather than pretend. Options include cluster randomization by account, region or store; switchback designs that alternate the whole system between variants over time windows; and interrupted time series with a control series. Each buys weaker causal claims than a user-level split, so name the weakness — contamination, temporal confounding, small cluster counts — rather than reporting the number as if it were an A/B result.
  • Why hold a canary for 48 hours instead of ramping as soon as it looks clean?
    Because many failures are time-dependent rather than volume-dependent: weekend traffic shape, nightly batches, cache warm-up, a scheduled dependency. An hour on a quiet afternoon exercises one load pattern. The hold is chosen to cover a real cycle — including a peak — so the ramp decision is based on the system's worst realistic conditions rather than its calmest.
  • Can canary data be pooled into the final experiment analysis?
    Sometimes, if the canary arm was genuinely randomized at the user level and the variant did not change between stages. But it is easy to get wrong: canary slices are often routed by region or client version rather than randomly, and ramping changes the assignment population mid-flight. Treat pooling as a design decision made in advance, not a convenience discovered when power is short.

saying these in an interview costs you the question

  • Concludes a change is better because a 1% canary showed no problems
  • Reads a fractional conversion effect off a tiny canary slice
  • Treats a region-routed or client-routed slice as a randomized arm
  • Ramps immediately on a clean hour without covering a peak load cycle
  • Runs a 50/50 split on a structural change with no prior exposure stage

context