A change passed every canary stage, reached 100% of the fleet, and the failure surfaced two days later. Which classes of failure does a progressive rollout structurally fail to catch, and what controls would you add for them?
answer
- time, scale, event, population
- 1% of load is 1% of the pressure
- leaks and disks need hours
- the batch job never ran during the bake
- keep the change attributable for days
basics
~20 sProgressive rollouts miss anything that needs time, scale, a specific event, or a specific population: slow leaks and disk fill, dependencies that only saturate at full traffic, monthly batches, and one tenant or old client. Each needs its own control.
solid answer
~1 minA canary tests a small slice of traffic for a short time, so four things escape it by construction. **Time**: leaks, unbounded caches, filling disks, log rotation, credential and certificate refresh — all incubate for hours or days. **Scale**: anything that only breaks when the *whole* fleet does it, because at 1% the shared database, connection pool, quota or hot cache key never gets near saturation, and a retry or fan-out pattern that is fine at 1% can be a stampede at 100%. **Events**: a monthly billing run, a weekly partner file, a quarter-end path that the canary window never touched. **Population**: a tenant, a locale, or a mobile client version that upgrades over weeks. The controls are different for each. Load-test the change at full expected volume before rolling. Keep the change behind a flag so deployment is not exposure. Slice SLIs by tenant and client version so a concentrated failure is visible. Trigger scheduled jobs synthetically rather than waiting. And keep the change attributable after full rollout — annotate deploys and hold a watch window at least as long as the longest incubation period you know about, because that two-day gap is exactly where attribution usually gets lost.
go deeper
Know that passing a canary is evidence, not proof, and be able to give one example of something a small short rollout cannot see — a slow memory leak or a monthly batch job.
Be ready to sort the misses into categories — needs time, needs scale, needs a specific event, needs a specific population — and explain why 1% of traffic applies only 1% of the pressure to a shared dependency.
Demonstrate the controls: score leading indicators like heap slope against a control, load-test at full volume before rolling, trigger scheduled jobs synthetically, slice SLIs by tenant and client version, and hold a watch window after the rollout completes.
Own the honesty of the release story. State what your rollout process does not verify, decide which risk classes require a control other than the canary, and make change attribution over a multi-day window a platform guarantee rather than an on-call memory exercise.
## The structural argument A progressive rollout makes one bet: that a small sample, observed briefly, predicts the whole. That bet is sound for failures that are *uniform in traffic and immediate in time* — a broken endpoint, a serialization bug, a latency regression on the hot path. It is unsound for everything else, and the categories are enumerable. ## Category 1 — needs time The defect exists from the first request but produces no visible symptom until an accumulating quantity crosses a threshold. - A leaked object, connection, thread or file handle; heap or file-descriptor exhaustion hours later. - An unbounded cache, or one whose eviction policy changed, that behaves perfectly while it is small. - A disk filling with logs, temp files or retained data, where the log line you added is 200 bytes per request and the volume is a million requests an hour. - Credential, token, lease or certificate refresh, where the code you changed runs once a day on the credential's schedule. **Controls.** Score leading indicators rather than outcomes — heap growth slope, file-descriptor count, disk-free trend, all compared against a same-age control, so a leak is visible as a divergence in minutes even though the crash is hours away. Run the change continuously in a soak environment. Hold a post-rollout watch window sized to the longest known incubation period. ## Category 2 — needs scale The defect is invisible at 1% because the pressure it creates is 1% of the pressure that breaks something. This is the category that most often turns a clean rollout into an outage minutes after the last stage. - A shared backend that saturates only under full load: connection pool exhausted, a database at its own concurrency limit, an API quota, a rate limiter that trips at aggregate volume. - A hot key or a cache stampede that only forms when the whole fleet misses simultaneously. - Fan-out amplification: one extra downstream call per request is trivial at 1% and can double a dependency's load at 100%. - Coordination effects — every instance refreshing config on the same schedule, thundering-herd reconnects. **Controls.** A load test that drives the *change* at full expected volume against a realistic dependency, before the rollout starts, is the only control that reliably compresses scale. Beyond that: watch the *dependency's* saturation signals during the rollout rather than only the service's own SLIs, and stage the last rungs at increments large enough to see the trend before the final jump — going 25% to 100% hides the knee that 25% to 50% would have shown. ## Category 3 — needs a specific event The changed path is only exercised on a schedule or by an external trigger: the nightly batch, the weekly partner file, month-end close, an annual renewal, a leap-day or daylight-saving boundary. **Controls.** Trigger the job synthetically against the canary instead of waiting for its schedule, or time a rollout stage to span it deliberately. Where neither is possible, record that the rollout did not verify that path, and treat the first real execution as a supervised event with someone watching. ## Category 4 — needs a specific population A single tenant's data shape, a locale, an old mobile or desktop client version that upgrades over weeks, a partner using an API in a way nobody documented. Random sampling misses concentrated populations, and for client-side software the "rollout" is an app-store or upgrade curve you do not control at all. **Controls.** Slice SLIs by tenant and by client version so a failure concentrated in one of them is visible instead of being averaged into a healthy aggregate. Deliberately include the risky cohorts in an early stage rather than sampling for them. ## The control that spans all four **Decouple deployment from exposure.** If the new behaviour sits behind a runtime switch, the binary can be fully rolled out while the behaviour is enabled separately and, crucially, disabled in seconds when a late failure appears — without a redeploy. That turns a two-day-late discovery from a rollout problem into a configuration change. It is not free: it adds a code path that must itself be tested in both states, and the switch has to be exercised occasionally rather than assumed to work. One asymmetry to keep in mind: some changes are not reversible by reverting the code at all, because the new version has already written data, published messages or migrated state that the old version cannot read. For those, the question of how you recover is its own discipline, and it should be answered before the rollout begins rather than during the incident. ## And the meta-control: attribution The reason a two-day-late failure is so expensive is rarely the failure itself — it is that nobody connects it to a rollout that finished on Tuesday. Annotate deployments on the dashboards the on-call actually looks at, keep the fleet's version state queryable, and make "what changed in the last N days" a question with an instant answer, where N is longer than your longest known incubation period. That single practice converts a mystery into a rollback candidate.
- Why can a change that adds one downstream call per request pass every canary stage and still take down a dependency?At 1% the extra call adds 1% of the eventual load, which the dependency absorbs invisibly. At 100% it can double that dependency's request rate, crossing its connection-pool, quota or concurrency limit. The canary measured the service's own SLIs, which looked fine. The control is to watch the dependency's saturation signals during the rollout and to load-test the change at full expected volume beforehand.
- How would you catch a memory leak during a canary rather than after the rollout?Score the leading indicator instead of the outcome. Compare the canary's heap growth slope, file-descriptor count and RSS trend against a same-age control; a leak shows as a divergence within minutes even though exhaustion is hours away. Waiting for an out-of-memory event requires a bake longer than the incubation period, which is usually longer than any acceptable rollout.
- A change touches code that only runs during the monthly billing batch. What does the rollout actually verify?Only that the new version serves normal traffic — the changed path was never executed. Either trigger the batch synthetically against the canary with representative input, or schedule a rollout stage to span a real run with someone watching. If neither is possible, record explicitly that the path is unverified and treat the first real execution as a supervised event rather than assuming the rollout covered it.
- Why is jumping a rollout from 25% straight to 100% riskier than it looks?Load-dependent effects are non-linear, and the interesting behaviour is at the knee. Going 25 to 50 to 100 gives you two observations of how the dependency's saturation signals respond to a doubling, so you can see a trend bending before you commit. A single jump to full traffic gives you the outcome with no chance to abort on the way.
saying these in an interview costs you the question
- Treats a passed canary as proof the change is safe
- Assumes any defect appears within the bake window
- Ignores that shared dependencies only saturate at full traffic
- Watches only the service's own SLIs, never the dependency's
- Cannot tell days later which rollout preceded a failure