An automated canary check scores the new version by comparing its error rate and latency against the metrics the rest of the production fleet reported over the same hour. Why is that comparison misleading, and what should the canary be compared against instead?
answer
- compare like with like
- the canary process started minutes ago
- cold caches and empty connection pools
- run a control of the old version
- same start time, same host class
basics
~20 sThe fleet has been running for days; the canary started minutes ago. Cold caches, unfilled connection pools and runtime warm-up make a healthy canary look slow. Compare it against a control running the old version, started at the same time on the same class of host.
solid answer
~60 sThe comparison is unfair in a way that has nothing to do with the code. A freshly started process has empty local caches, an unfilled connection pool, an unwarmed runtime, and often a cold downstream cache keyed by its own identity, so its first minutes of p99 latency look terrible even if the change is a no-op. Meanwhile the long-running fleet has been serving for days. Comparing yesterday's numbers is worse still, because it also folds in a different traffic mix, a different batch schedule and a different set of neighbours. What you want is a *control*: an equally fresh deployment of the **old** version, launched at the same time, the same size, on the same host class and in the same zones, receiving comparably routed traffic. Then the only variable left between the two groups is the change. This is why Spinnaker's Kayenta and similar canary-analysis systems deploy a baseline alongside the canary rather than scoring against production. Whatever you compare against, exclude the warm-up window from the scoring rather than pretending it is signal.
go deeper
Know that a canary's numbers only mean something next to a comparison, and that a brand-new process is slower at first because its caches and connection pools are empty.
Be ready to name the specific warm-up effects — cold caches, unfilled pools, an uncompiled runtime — and to explain why a control deployed at the same time on the same hardware removes them from the comparison.
Show that you treat the rollout as a controlled experiment: matched age, host class and zones; relative rather than absolute thresholds; and missing data treated as inconclusive rather than as a pass.
Own the cost of the control. Running a baseline for every stage of every rollout is real capacity and real pipeline time, so be able to say which risk classes justify it, which get the cheaper matched-subset comparison, and how you keep the weaker gate from being quietly used everywhere.
## The comparison is the experiment A canary is an experiment with one treatment group and one control group, and every rule about confounded experiments applies. If the treatment group differs from the control in more ways than the change you are testing, the result measures the confounder. In a rollout the dominant confounder is not the code — it is **age**. ## What is different about a process that started two minutes ago - **Local caches are empty.** An in-process cache with a 95% steady-state hit rate starts at 0%, so every miss becomes a downstream call. Tail latency is dominated by misses. - **Connection pools are unfilled.** The first requests pay TCP and TLS handshakes, and possibly a DNS lookup, that a warm instance does not. - **The runtime is cold.** On a JIT-compiled runtime, hot methods are still interpreted; on any runtime, the page cache has not been populated and lazily initialised code paths are being touched for the first time. - **Downstream state is cold.** Prepared statements, session tokens, authorisation caches and even the database's own plan cache for that client may need to be established. - **Autoscaling has not settled.** A new instance may briefly hold more or fewer connections than its share, depending on how the load balancer's least-connections or round-robin policy treats a new member. None of that is a property of the change. All of it moves error rate and p99 latency in the bad direction for the first seconds to minutes. ## The three comparisons and what each is worth **Against the running fleet.** Confounds age, and also confounds host generation: if the fleet spans two instance types and the canary happens to land on the older one, the canary loses on latency for a reason nobody shipped. Usable only if you exclude the warm-up window *and* restrict the comparison to a matched subset — which is most of the work of running a baseline anyway. **Against a historical window (yesterday, last week).** Confounds age plus traffic mix, plus whatever else was different: a batch job that runs on Tuesdays, a marketing campaign, a partner's retry storm, a different weather of neighbours on shared hardware. It is the weakest comparison and the most common, because the data is already there. **Against a purpose-deployed baseline.** Deploy the *old* version fresh, at the same moment, at the same replica count, on the same host class, in the same zones, and route it a comparable slice of traffic. Now canary and baseline share their age, their hardware, their neighbours and their traffic window, and the difference between them is the change. This is the design used by automated canary analysis systems such as Spinnaker's Kayenta, and it is why they talk about *canary versus baseline* rather than *canary versus production*. ```text stable fleet (old version, days old) <- not the control baseline (old version, 2 min old) <- the control canary (new version, 2 min old) <- the treatment ``` ## What still breaks a properly built comparison Running a baseline is necessary, not sufficient. - **Traffic is not actually comparable.** If assignment is sticky by tenant, one group may catch a heavy tenant, and then you are comparing different workloads. Check the realised request mix, not the configured split. - **The window is too short.** Two minutes of data on a low-volume endpoint has no power to detect anything; the result is noise dressed as a verdict. - **The metric is non-stationary.** Deploying across the morning ramp means both groups are changing under you; that is fine for a *relative* comparison and fatal for an absolute threshold. - **Only one of the two is scraped correctly.** A canary whose metrics are missing scores as "no errors", which is the most dangerous false pass there is. Treat missing data as inconclusive, never as healthy. ## The cost, and when to skip it A baseline costs real capacity — you are running a third group for the length of the rollout — and it doubles the deployment work for every stage. For a low-risk service with plenty of traffic, excluding the first few minutes and comparing the canary against a matched slice of the fleet is a defensible shortcut. For a change whose blast radius is the whole product, the extra instances are cheap next to the outage you are trying to avoid. The judgement call is what a wrong verdict costs, in both directions: a false pass ships a bad change, and a false fail teaches people to ignore the gate.
- If a baseline deployment costs capacity you do not have, what is the cheapest honest alternative?Exclude the warm-up window from scoring and restrict the comparison to a matched subset of the fleet — same instance type, same zone, same time window — rather than the fleet average or a historical baseline. You still carry the age confounder for anything that warms slowly, so pair it with a relative-change threshold rather than an absolute one, and be explicit that the gate is weaker.
- Your canary's metrics stop being scraped halfway through the window. What should the analysis conclude?Inconclusive, not healthy. Absent data reads as zero errors in most scoring implementations, which is the most dangerous possible false pass. The gate should require a minimum number of observed samples per scored metric and fail closed — hold the rollout and escalate — when the samples are not there.
- Why does a rollout during the morning traffic ramp break absolute thresholds but not relative comparison?Absolute thresholds assume the metric is stationary: "p99 under 300 ms" is false at peak and trivially true at 3 am, regardless of the change. A relative comparison against a control experiencing the same ramp cancels the trend out, because both groups move together and only the difference between them is scored.
saying these in an interview costs you the question
- Compares a two-minute-old canary against a week-old fleet
- Uses yesterday's metrics as the control
- Scores absolute thresholds instead of the difference from a control
- Treats missing canary metrics as a clean result
- Ignores that canary and baseline may receive different traffic mixes