A canary deploy caught a fatal bug before any customer saw an error. Do you analyze it like an incident, and what would that analysis look at?
answer
- all the information, none of the impact
- review by margin, not by outcome
- designed control or lucky sample?
- would it catch a slower variant?
- saves without review let the margin drift
basics
~20 sYes, when the margin was thin or the save was luck rather than design. The analysis asks how close it came, whether the control that caught it was intentional and repeatable, and what would have happened had it fired ten minutes later.
solid answer
~50 sNear misses are the cheapest incidents you will ever get — full information, no customer impact, no time pressure — and most organizations throw them away because nothing broke. My criterion for reviewing one is margin, not outcome: if the control that saved us was designed for this and had comfortable headroom, a line in the deploy log is enough. If it caught the bug by accident, or barely, or because one person happened to be watching the right graph, it gets a review. The analysis asks a different question from an incident postmortem: not "why did this hurt?" but "why didn't it?" Was the canary population large enough to detect this reliably, or did we get a lucky sample? Would the automated analysis have failed it, or did a human notice? How much longer would the bake have needed for a slower-manifesting bug? The finding is usually about the control's reliability, and that is worth knowing before the day it doesn't hold.
go deeper
Know that a near miss is an event that could have caused an outage but didn't, and that it is worth writing down because it shows a real weakness without any customer cost.
Explain the margin-based selection criterion and be able to distinguish a control that was designed to catch the problem from an accidental save that will not repeat.
Quantify the closeness and the counterfactual impact, and show how the detection power of a canary depends on exposed population and bake duration for a given failure rate.
Own the drift problem: repeated unreviewed saves let the safe operating margin erode invisibly, so the org needs a cheap standing way to surface near misses and a proportionate review that does not consume incident-review capacity.
## Why near misses are worth the time A near miss is an event where the conditions for a serious incident were present and something stopped it. In an outage you pay for the information with customer impact. In a near miss you get most of the same information free: the defect existed, the path to production was real, the failure mode is observable — only the impact is missing. Aviation and medicine formalized near-miss reporting decades ago for exactly this asymmetry, and software teams mostly have not, because the trigger for a review is usually "did the SLO burn?" rather than "how close was it?" The practical consequence is that organizations learn about their controls only when the controls fail. ## The selection criterion is margin, not outcome Reviewing every save would be its own kind of toil, and review capacity competes directly with real incidents. The useful filter is how much room there was: - **Designed control, wide margin.** The canary failed automated analysis at 2% of traffic within four minutes, exactly as intended. Log it, move on. This is the system working. - **Designed control, thin margin.** It failed analysis at minute 29 of a 30-minute bake. The control held, but a slightly slower manifestation walks straight past it. - **Undesigned save.** Someone happened to be looking at a dashboard. The canary sample happened to include the one shard with the affected data. A dependency happened to be slow, which throttled the rollout. None of these repeat on demand — and treating them as evidence the system is safe is how a fluke gets recorded as a control. An undesigned save should be reviewed with the same seriousness as a real outage, because the next occurrence has no protection at all. ## What the analysis asks The questions invert an incident postmortem: 1. **How close was it?** Quantify. What fraction of traffic was exposed, for how long, and how far from the failing threshold did the SLI sit? "It was fine" is not an answer; "errors reached 0.4% against a 0.5% abort threshold" is. 2. **What actually caught it — and was that its job?** Distinguish the control that was designed to catch this class of thing from whatever incidentally did. 3. **Would it have caught a variant?** Slower onset, a bug affecting 1 in 500 requests instead of 1 in 5, a failure only visible in a region the canary does not cover. Canary detection power depends on the exposed population and the bake duration; a bug rarer than your canary's sample size will not show up. 4. **What was the counterfactual impact?** If this had reached full production, what would it have cost — how many users, how much of the error budget? This is the number that funds any follow-up work, and without it a near-miss review produces findings nobody prioritizes. 5. **Would the next line of defense have held?** If the canary had passed it, would monitoring have paged quickly, and was the rollback path ready? ## Where near misses come from Beyond canaries: an alert that fired and self-resolved before anyone acknowledged it, a failover that worked but took far longer than expected, a disk that reached 92% before a cleanup job ran, a retry storm that stayed just under the point where the database would have tipped, a deploy rolled back on a hunch. Each carries the shape of a real incident minus the impact. A cheap way to surface them systematically is to ask, at the end of any routine review, whether anything nearly went wrong this week — because near misses have no alert of their own and nobody is paged when disaster does not happen. ## The failure mode to name Repeated near misses with no review are how a margin quietly disappears. Each save is read as proof the system is resilient, the acceptable operating point drifts toward the edge, and the population that would have flagged the drift is exactly the reviews nobody ran. The point of counting near misses is to notice the drift while it is still free to correct. ## In an interview Give a concrete near miss, state the margin as a number, and say what you changed as a result — the canary population, the bake duration, an alert that should have fired first. Saying "we would write a postmortem" is a policy statement; saying "errors peaked at 0.4% against a 0.5% abort threshold, so we shortened the analysis window" is an answer.
- How do you stop near-miss reviews from becoming their own source of toil?By gating on margin and by making the review proportionate. A wide-margin save logged automatically costs nothing. A thin-margin or undesigned save gets a short written review focused on the control, not a full incident document. If the volume is still high, that is itself the finding — a system producing many near misses a week has a systemic problem worth one investigation rather than twenty reviews.
- Where should the counterfactual impact estimate come from?From the observed behaviour in the exposed population, extrapolated to full traffic. If 2% of traffic saw a 30% error rate for six minutes, you can state what 100% for the same duration would have cost in failed requests and in error-budget minutes. Keep it as a range and label it an estimate — its job is to size the follow-up work, not to be precise.
- Does a near-miss review belong in the same repository as incident postmortems?Yes, tagged as a near miss. Keeping them together is what makes patterns visible later — three near misses on the same dependency is exactly the signal aggregate analysis is looking for, and it disappears if they live in a side channel. The tag preserves the distinction so incident counts and duration statistics are not distorted.
saying these in an interview costs you the question
- Reviews only events that breached the SLO
- Counts a lucky catch as proof the control works
- Cannot state how close the event came in numbers
- Assumes a canary detects any bug regardless of its rate
- Treats every near miss as requiring a full postmortem