Three outages this quarter had different proximate causes, but all three involved the same shared configuration service. How do you approach the analysis across them?
answer
- the unit of analysis is the set
- the cause differs, the factor repeats
- normalize for exposure before blaming
- involvement is not contribution
- aggregate impact is what funds structural work
basics
~20 sAnalyze the set, not the incidents. Tag contributing factors across postmortems, find the factor that repeats, normalize for how much traffic that component actually carries, then argue a structural investment using aggregate impact rather than three separate small fixes.
solid answer
~60 sPer-incident analysis cannot see this pattern, because each postmortem was individually correct — three different triggers, three reasonable fixes. The pattern only appears when you analyze the corpus, so I tag contributing factors consistently across postmortems and look for the factor that repeats rather than the cause that differs. Before concluding anything I check the base rate: a component every request touches will appear in most incidents simply because it is on every path, so I normalize by exposure — incidents per unit of traffic or per dependent service — otherwise I will fund the wrong fix. If it survives that check, the decision is genuinely contested: keep shipping cheap per-incident fixes, or fund something structural like isolating the dependency, adding a cached fail-static path, or splitting the service by criticality tier. The argument for the structural option is made in aggregate — total impact minutes, error budget consumed across all three, engineering hours spent responding — because no single incident justifies a quarter of work, and three of them together often do.
go deeper
Know that patterns across several postmortems can matter more than any single incident, and that the same component turning up repeatedly is worth flagging even when each outage had a different cause.
Explain how consistent tagging of contributing factors and comparable impact numbers make a set of postmortems analyzable at all, and what you would look for across them.
Show the base-rate correction — normalizing by exposure and separating involvement from contribution — and pick a bounded blast-radius fix you could ship before anyone funds a rewrite.
Own the decision and the mechanism behind it: aggregate the impact into a fundable argument, choose between hardening and decoupling with the costs stated, and name who holds the cross-incident view when the pattern crosses team boundaries.
## The blind spot of per-incident analysis A postmortem is scoped to one event, and that scoping is deliberate: it keeps the analysis concrete and the findings actionable. It also guarantees that a class of problem is invisible. If a shared component contributes a little to many incidents but is the headline cause of none, every individual postmortem will correctly attribute the outage to something else, and the pattern will be discussed only in hallway conversations where nobody can fund anything. Seeing it requires a second, slower analysis whose unit is the set of postmortems rather than the incident. ## Making the corpus analyzable This only works if factors were recorded in a comparable form, which is a standard you have to set before you need it: - **Consistent tagging.** Contributing factors carry tags — the affected service, the class of factor (rollout policy, alert coverage, dependency behaviour, capacity), and the incident phase. Free-text-only postmortems cannot be aggregated without re-reading them all. - **Comparable durations.** Time-to-detect and time-to-mitigate defined the same way across incidents, so they can be summed and compared across quarters. - **Impact recorded in one currency.** Usually error-budget minutes or affected-user-minutes, so three incidents can be added together into a number that means something. Then the quarterly question is not "what caused each of these?" but "which factor appears in the most incidents, and which factor accounts for the most impact?" Those are often different components, and the second one is usually the better investment. ## The base rate trap The most common error here is confusing exposure with fault. A configuration service that every request path touches will appear in a large share of incident timelines by construction — it is on the path, so it is in the story. Two corrections: 1. **Normalize by exposure.** Incidents per dependent service, or per million requests served through it, rather than raw counts. A component involved in three of twelve incidents while serving every request may be more reliable than one involved in two of twelve while serving a tenth of traffic. 2. **Distinguish involvement from contribution.** Being present in the timeline is not the same as being a contributing factor. The test is the materiality one applied per incident: would a different behaviour from this component have prevented or reduced the impact? A component that merely propagated someone else's failure faithfully is not the finding — though "we have no fallback when it is unavailable" might be. ## The decision, and what it costs Assume the pattern survives. Now the choice is real: | Option | Cost | Risk | |---|---|---| | Continue per-incident fixes | Days each, already budgeted | The pattern continues; you pay repeatedly and learn nothing new | | Reduce blast radius (fail-static cache, last-known-good, per-tier isolation) | Weeks | Often the best value; bounded work, breaks the shared-fate link | | Harden the component itself | Weeks to months | Fixes the source but the dependency remains a single point of failure | | Remove the shared dependency from the critical path | A quarter or more | Highest cost, hardest to fund, sometimes the only real answer | The aggregate framing is what makes the middle options fundable. One incident of eleven minutes justifies nothing structural. Three incidents totalling ninety minutes of impact, a large fraction of the quarter's error budget, plus the on-call hours and the review time, is an argument a business can act on. Where a budget policy exists, exhausted budget attributable to one dependency is the strongest lever available, because the consequence is already agreed in advance rather than negotiated after the fact. ## Being honest about the limits Three data points is a weak sample and you should say so. Deciding to fund structural work on three incidents is a judgment call about downside risk, not a statistical conclusion, and it is more defensible when there is a mechanism that explains the pattern — a genuine shared-fate coupling — rather than only a correlation in the tags. Look for corroborating evidence too: near misses involving the same component, and pages that never became incidents. ## The organizational half The pattern often crosses team boundaries, which is exactly why nobody owns it: each team's own postmortem quality is fine. Somebody has to hold the cross-incident view — a periodic reliability review, a designated owner of the postmortem corpus — and be able to fund work outside a single team's roadmap. Naming that gap is usually a stronger interview answer than any particular technical fix, because it is the reason the pattern persisted for three incidents in the first place.
- How would you tell a genuine shared-fate pattern from a coincidence of three unrelated bugs?Look for a mechanism, not just a count. A real pattern has an explanation you can state — every caller reads config synchronously on the request path with no cached fallback, so any unavailability is immediate user impact. If the only thing the three share is the component's name in the timeline, and each failed through a different mechanism it handled correctly, that is exposure, not shared fate.
- What do you do when the structural fix is clearly right but nobody will fund it?Convert it from an engineering opinion into a stated risk with a number attached: aggregate impact so far, projected recurrence, and what the next occurrence costs. Then take the bounded step you can fund — a fail-static fallback or isolating the highest-criticality callers — which reduces blast radius without the full investment. If a budget policy exists, an exhausted budget traced to this dependency is the cleanest forcing function available.
- How often should this cross-incident analysis run?Frequently enough that patterns surface while they are still cheap, and rarely enough that it stays a real analysis rather than a status meeting — quarterly suits most teams, monthly for high-incident-rate services. The important property is that it has a named owner and an audience able to fund work, otherwise it produces observations that go nowhere.
saying these in an interview costs you the question
- Blames the most-touched component without normalizing for exposure
- Treats each postmortem as complete because each was individually correct
- Argues for a rewrite from one incident's impact
- Confuses appearing in a timeline with contributing to the outage
- Assumes free-text postmortems can be aggregated without a tagging standard