You roll a change out to 1% of production traffic, the canary's error rate and latency stay clean for an hour, and then the change breaks one enterprise customer's integration the moment it goes to 100%. Why can a 1% traffic canary contain 0% of the users a change affects, and how would you choose the canary population instead?
answer
- think about who, not how many
- rare paths contribute very few requests
- one tenant may send zero canary requests
- 0.99^20 is about 82 percent
- select the cohort, then sample randomly
basics
~20 sA traffic percentage is a random sample, not a representative one. A rare code path, a single tenant, or one old client version can send so few requests that the canary receives none of them. Choose the canary population deliberately, not just its size.
solid answer
~60 sRouting 1% of requests picks requests at random, and randomness is only representative when the thing you are testing is spread evenly across requests. If a customer sends twenty requests in the canary window, the chance none of them lands on the canary is `0.99^20`, about 82% — the canary is genuinely blind to that customer. The same holds for a code path that is 0.5% of traffic, a locale, or a mobile client version that only a slow-upgrading minority still runs. So I size the canary by how many events of the *risky path* it needs to see, not by a round percentage, and where the risk is concentrated I select the population instead of sampling it: internal and dogfood traffic first, then a named cohort of tolerant tenants, then random traffic, and the highest-volume or highest-value customers last. For anything stateful I hash on user or tenant ID so a session does not flip versions mid-flow, and I make the assignment a label on the telemetry so I can actually attribute an error to the canary.
code
python · 10 linescanary_share = 0.01
requests_from_customer = 20
p_never_hits_canary = (1 - canary_share) ** requests_from_customer
print(round(p_never_hits_canary, 3))
rps = 1000
path_share = 0.005 # the changed code path's share of traffic
bake_seconds = 15 * 60
samples = rps * canary_share * path_share * bake_seconds
print(round(samples))go deeper
Know that a canary sends a small share of real production traffic to the new version before everyone gets it, and be able to say why that is safer than switching everything at once.
Be ready to explain that a traffic percentage is a random sample and to do the arithmetic out loud: how many requests a 1% canary gives the specific code path you changed, and how likely it is to miss a low-volume customer entirely.
Show that you design the population, not just the percentage — internal traffic, then a chosen cohort, then random share, with the highest-volume customers staged last — and that you label telemetry by version so canary errors are attributable at all.
Own the tradeoff between early signal and representative signal: deliberately chosen cohorts find concentrated risks fast but are unrepresentative, while random shares are representative but blind to the tail. Be able to say what your release policy requires for each risk class and who maintains the cohorts.
## The claim a percentage rollout is really making "1% of traffic is on the new version" is a statement about request routing. The statement people *hear* is "we have tested the change against 1% of reality". Those are only the same claim when the failure you are looking for is spread uniformly across requests — a change that breaks every request, or slows down the hottest endpoint. For anything narrower, a percentage of traffic is a sample, and a small random sample of a skewed population misses the tail almost every time. ## The arithmetic, because it is the whole point Two numbers decide whether a canary can see a thing. **Coverage of a cohort.** If a customer sends `n` requests during the canary window and each request has probability `p` of being routed to the canary, the chance the canary never sees that customer is `(1 - p)^n`. At `p = 1%` and `n = 20`, that is `0.99^20 ≈ 0.82`. Four times out of five you learn nothing about them. To be reasonably confident of at least one hit you need `n` in the hundreds at 1%. **Sample count for a rare path.** Take a service at 1,000 requests per second. A 1% canary receives 10 rps. If the changed code path is 0.5% of traffic, the canary exercises it 0.05 times per second — three times a minute. A fifteen-minute bake gives about 45 executions, which cannot distinguish a jump from a 0.1% error rate to a 1% error rate from noise. The canary is green because it is empty, not because it is healthy. ## Choosing the population instead of sampling it The fix is to treat "who is in the canary" as a design decision with its own ladder, roughly: 1. **Synthetic and internal traffic.** Cheap, fully attributable, and you own the complaint. It catches gross breakage before any customer is exposed. 2. **Dogfood / employee accounts.** Real workflows, real data shapes, a population that tolerates breakage and reports it in words rather than in a support ticket. 3. **A named cohort.** Tenants or accounts chosen because they exercise the changed path, or because their tolerance is known. This is what you use when the risk is concentrated — the integration partner whose contract you touched should be *in* the canary, deliberately, not sampled into it with probability 0.18. 4. **Random traffic share.** Now the percentage does its real job: measuring aggregate SLIs on a representative slice. 5. **The whale last.** If one customer is 70% of your volume, any canary containing them is not a canary, and any canary excluding them tells you nothing about them. Stage them as their own step, with their own bake and their own watchers. ## Per-request versus per-entity assignment Random per-request routing is fine for a stateless read. It is actively harmful for a multi-step flow: a checkout whose first call hits the new version and whose second call hits the old one exercises a version combination you never intended to test, and the resulting error is attributable to neither. Hash on a stable key — user ID, tenant ID, session — so an entity stays on one version for the whole rollout. That is also what makes the analysis honest: you can count *users* affected, not just requests. ```text canary_bucket = hash(tenant_id) % 1000 < 10 # sticky 1% by tenant, not by request ``` The cost of sticky assignment is that your 1% of *entities* is not 1% of *traffic* — if the hash happens to catch a heavy tenant, the canary may carry 8% of requests. Measure the realised share rather than assuming the configured one. ## Attribution None of this works if you cannot tell canary telemetry from baseline telemetry. Every request the canary serves should carry the version and the assignment as a label so error rate, latency and business metrics can be sliced by it. Without that label a 1% canary contributes 1% of the errors to an aggregate dashboard, where a total failure of the new version looks like a 1% blip and is invisible against normal noise. ## What this costs Deliberate populations are more work than a percentage: someone maintains the cohort, cohorts go stale, and a cohort of tolerant customers is by definition unrepresentative of the intolerant ones. The honest position is that population selection buys you *early* signal about specific risks, and the random share buys you *representative* signal about aggregate ones. A serious rollout uses both, in that order.
- If you assign the canary by hashing the tenant ID rather than per request, what new problem do you have to watch for?The realised traffic share stops matching the configured entity share. Hashing 1% of tenants into the canary can send it 8% of requests if a heavy tenant lands in the bucket, or almost none if only idle tenants do. Measure the actual request share the canary receives and size your capacity and your analysis from that, not from the configured percentage.
- One customer is 70% of your traffic. How do you canary at all?Accept that they are their own rollout stage. Everything else — internal, dogfood, the long tail of small tenants — goes first and validates the change generally. Then the whale gets a dedicated stage with its own bake, its own watchers and its own abort criteria, ideally starting with a fraction of their traffic if their client supports it. Treating them as 70% of a random sample means every canary is either blind or full-blast.
- How do you keep a canary cohort from going stale?Define it by a property rather than by a list of account IDs — tenants on a given plan, accounts that called the endpoint you are changing in the last week, or employees. A hand-maintained list drifts until it contains customers who churned and misses the integration that was built last quarter. Re-derive the cohort at rollout time and log which entities it resolved to.
saying these in an interview costs you the question
- Thinks 1% of traffic means 1% of users are affected
- Assumes a green canary proves the change is safe
- Sizes the canary by cost or by a round number only
- Routes per request even for multi-step stateful flows
- Cannot slice telemetry by version, so canary errors vanish in the aggregate