skip to content

You are setting the bake time for each stage of a canary rollout — how long the new version soaks at 1% before it advances to 10%. How do you decide that number, and why is a 10-minute bake worthless for some changes?

level: seniorimportance: should knowfreq 50%

answer

  1. two clocks: statistics and incubation
  2. enough events to detect the delta
  3. how long until the feared failure appears
  4. 3 am is not peak traffic
  5. raising share buys evidence faster than waiting

basics

~20 s

Two clocks set the bake: how long it takes to collect enough events at that traffic share to detect the regression you care about, and how long the failure mode you fear needs to appear. Bake for the longer of the two.

solid answer

~1 min

I decide bake time from two independent constraints and take whichever is longer. The first is statistical: at this stage's traffic share, how long until the changed code path has produced enough events to distinguish a real regression from noise? At 1% of a 1,000 rps service, a path used by 0.5% of requests produces about three events a minute, so a ten-minute bake gives thirty samples and proves nothing. The second is temporal: what is the incubation period of the failure I am afraid of? A memory leak that OOMs after six hours, a connection pool that only exhausts under the hourly batch, a cache that fills over a day, a certificate refreshed nightly — none of them appear in ten minutes at any traffic share. So the bake must cross at least one full period of the fear. Against that, long bakes cost real things: a longer mixed-version window, emergency fixes queued behind the rollout, and rollouts overlapping each other. My usual shape is short automated bakes at the tiny stages where the check is a fast smoke signal, and one deliberately long soak at a meaningful percentage — often overnight, so it crosses a diurnal cycle and the nightly batch — before the fleet goes wide.

code

python · 10 lines
python
total_rps = 1000
canary_share = 0.01
changed_path_share = 0.005

for bake_minutes in (10, 60, 240):
    events = total_rps * canary_share * changed_path_share * bake_minutes * 60
    print(f"{bake_minutes:>4} min at 1%: {events:>6.0f} events on the changed path")

# same evidence, less wall clock: raise the share instead
print(round(total_rps * 0.10 * changed_path_share * 10 * 60), "events in 10 min at 10%")

go deeper

for a junior

Know that a canary stage is held for a period before advancing, and that the point of the wait is to collect enough real production evidence to judge the new version.

for a middle

Be able to compute the evidence a stage actually produces — traffic rate times canary share times the changed path's share — and explain why a small share for a short time proves nothing.

for a senior

Show the second clock: name failure modes with incubation periods (leaks, scheduled jobs, credential refresh, disk fill) and say that the bake must cross one full period, or the control must change to a load test or a post-rollout watch window.

for a principal

Own the cost side. Long bakes lengthen the mixed-version window, queue emergency fixes and push teams toward batching. Be ready to state the rollout duration your organisation can afford, the express path for urgent changes, and who decides when a stage is skipped.

## Two clocks, not one People pick bake times the way they pick timeouts: a round number that felt reasonable once. There are actually two independent questions, and the bake must satisfy both. ### Clock one: statistical power A canary stage is a measurement, and a measurement needs samples. The relevant rate is not the service's request rate, it is the rate at which the *changed behaviour* is exercised inside the canary: ```text events_per_second = total_rps x canary_share x changed_path_share ``` At 1,000 rps, a 1% canary and a path that is 0.5% of traffic, that is 0.05 events per second — three per minute. A ten-minute bake yields thirty events. If the baseline error rate on that path is 0.1% and the regression takes it to 1%, thirty samples will usually show zero failures in both groups. The gate passes because the experiment had no power, not because the code is good. The consequence is a lever people forget: **raising the traffic share buys the same evidence faster than extending the bake**. Ten minutes at 10% gathers the same samples as one hundred minutes at 1%. Which one you choose is a blast-radius decision — more users exposed for less time, or fewer users for longer — and it should be made explicitly rather than by defaulting to a long bake at a tiny share. ### Clock two: the incubation period of the failure you fear Some failures are not sample-limited at all; they are time-limited. No traffic share makes them appear sooner, or the relationship is so weak that it may as well not exist: - **Resource accumulation** — a leaked object, an unbounded cache, a file handle never closed, a disk filling with logs. These surface when a threshold is crossed, hours in. - **Scheduled work** — an hourly reconciliation, a nightly batch, a weekly report, a monthly billing run. If the change touches that path, a bake that does not span the schedule tests nothing about it. - **Warm-state effects that only degrade later** — a cache whose eviction policy changed behaves fine while it is small. - **Expiry** — a token, a session, a lease or a certificate whose refresh path is the code you changed. The first refresh is the test, and it happens on the credential's clock, not yours. - **Diurnal position.** A bake that runs at 03:00 never sees peak concurrency. It tells you the version starts and serves; it says nothing about behaviour at the top of the day. The rule that survives contact with production: **bake at least one full period of the failure mode you are actually worried about**, and if that period is longer than any acceptable rollout, stop relying on the bake and use a different control — a load test at full volume, a synthetic trigger of the scheduled job, or an explicit watch window after the rollout with the change still attributable. ## What long bakes cost This is where judgement lives, because the costs are real and not always visible to the person choosing the number. - **A long mixed-version window.** Two versions serve simultaneously for the whole rollout. Every compatibility constraint — wire formats, cache entry shapes, shared state — must hold for that entire duration, and the longer the window, the more likely a rare cross-version interaction is exercised. - **Emergency changes queue behind it.** If a rollout occupies the pipeline for eight hours, the fix for an unrelated production bug waits, unless you have an express path that skips stages deliberately. Design that path *before* you need it, and require someone to own the decision to use it. - **Rollouts overlap.** If a rollout takes longer than your deploy interval, several are in flight at once. Attribution degrades: when something breaks, more than one change is mid-rollout. - **Batching pressure.** Teams facing a day-long rollout stop shipping small changes, and larger changesets are exactly what progressive delivery is trying to avoid. ## A shape that usually works Short, automated bakes at the smallest stages, sized by how long the automated check needs to gather its minimum samples — often minutes. Then a meaningful stage, say 10–25%, held long enough to cross a diurnal cycle and any scheduled job the change could touch — in practice, overnight. Then the fleet. The tiny stages catch gross breakage cheaply; the long soak catches the slow things; and only one stage in the sequence is paying the cost of a long hold. And say the quiet part in the design review: the bake time you can afford is a statement about how much risk you are accepting, not a technical constant. If the number was chosen because the pipeline felt slow, that is a decision about risk made by nobody.

  • You need more evidence from a canary stage. When is raising the traffic share better than extending the bake?
    When the failure is sample-limited rather than time-limited. Ten minutes at 10% gathers the same evidence as one hundred minutes at 1%, and it keeps the rollout short, so the mixed-version window shrinks and fixes do not queue. Extend the bake instead when the failure needs wall-clock time — a leak, a scheduled job, a credential refresh — because no traffic share makes those arrive sooner.
  • What do you do when the failure you fear has an incubation period longer than any acceptable rollout?
    Stop asking the bake to catch it. Use a control that compresses time or scale instead: a load test that drives the change at full volume, a synthetic trigger of the scheduled job, or a soak environment running the change continuously. Then keep the change attributable after the rollout — mark the deploy in dashboards and hold a watch window at least as long as the incubation period.
  • How do you keep an urgent fix from queueing behind a multi-hour rollout?
    Define an express path in advance: a documented set of stages that may be skipped, who is allowed to authorise it, and what minimum verification still runs. Ad-hoc bypassing under pressure is how a rollout control becomes theatre. The express path should also be exercised occasionally, because a route nobody has used is not a route you can rely on during an incident.

saying these in an interview costs you the question

  • Picks a round bake time with no reference to event volume
  • Believes any bake length can surface a six-hour memory leak
  • Bakes overnight only, so the change never sees peak traffic
  • Ignores that a long rollout leaves two versions serving for hours
  • Never considers raising the traffic share instead of waiting longer

context