A draft postmortem lists fourteen contributing factors for one outage. How do you decide which ones belong in the final analysis and which are noise?
answer
- everything noticed is not everything that mattered
- counterfactual on system state, not perception
- would it fit any other postmortem?
- who has a lever within a quarter
- an empty detection row means you stopped early
basics
~20 sKeep a factor if changing it would plausibly have prevented or materially reduced the impact, it can recur beyond this incident, and someone has a lever on it. Drop factors visible only in hindsight and factors so generic they would fit any postmortem.
solid answer
~50 sFourteen factors usually means the team confused "everything we noticed" with "what mattered." I apply three filters. First, materiality: if this factor had been different, would impact have been prevented or meaningfully reduced? That test must be applied to system state — the config was pushed globally, the alert threshold was ten minutes of averaging — not to what a person could have noticed, because that version is unfalsifiable and produces findings nobody can act on. Second, generality: does this factor apply beyond this one incident? A gap in rollout policy will show up again; a one-off typo will not. Third, tractability: does anyone in this organization have a lever on it within a quarter? Then I group survivors by phase — what let it ship, what delayed detection, what delayed mitigation, what amplified impact — and if a phase is empty I go back, because an empty detection row almost always means we stopped analyzing early rather than that detection was perfect.
go deeper
Know that a postmortem lists several contributing factors, and that not every observation from the incident becomes one — the useful factors are things that could be changed to prevent or shorten a future outage.
Explain the materiality test and why factors phrased as "if someone had noticed" are weaker than factors phrased as properties of the system.
Demonstrate live pruning with reasons, and use the phase grouping to show that a factor list concentrated on the defect leaves detection and mitigation time untouched.
Own the standard across the org: how specific a factor must be to count, how factors owned by other teams are escalated rather than dropped, and how the total volume of findings is kept within what the organization can actually fund.
## The problem with a long list A fourteen-factor postmortem is not more rigorous than a five-factor one; it is usually less. Long lists are produced by writing down everything anyone said during the review, and they have three costs. They dilute the findings that matter, they make the document unreadable so nobody reads it later, and they generate more follow-up work than the team can fund, which trains everyone that postmortem outputs are optional. Pruning is therefore part of the analysis, not editorial tidying. But pruning is also where bias enters, so the criteria need to be explicit. ## Filter 1: materiality Ask: **if this factor had been different, would the impact have been prevented or materially reduced?** This is a counterfactual, and counterfactuals are a sharp tool held the wrong way round most of the time. The version that works points at system properties: - "If the config had rolled out region by region, impact would have been one region for four minutes instead of global for forty." — keep, testable, actionable. - "If the alert had used a two-minute window instead of a ten-minute average, detection would have been eight minutes earlier." — keep. The version that does not work points at what a person could have perceived: - "If the reviewer had noticed the missing validation, the change would not have merged." - "If the on-call had checked the queue dashboard first, mitigation would have been faster." These read like findings but are hindsight in disguise: they are only obvious now that the answer is known, they cannot be falsified, and the only fix they imply is that people should be more careful, which is not a control. If a factor of that shape survives, convert it into the system property behind it: what made the missing validation invisible in the diff, or what made the queue dashboard the fourth place anyone looked. ## Filter 2: generality Does the factor exist outside this incident? Rollout policy, alert design, retry configuration, dependency fan-out, access to the rollback tool, runbook accuracy — all general, all will be involved again. A specific transposed digit in one constant is not general and belongs in the timeline as an event, not in the factor list as a finding. This filter also catches its own overcorrection. "Insufficient testing" and "inadequate monitoring" are so general they carry no information — they would fit any postmortem in the company. If you could paste a factor into an unrelated incident report without editing it, it is a category, not a finding. Push it down one level of specificity: not "insufficient testing" but "config schema is validated at runtime only, so no test in CI can reject a bad field." ## Filter 3: tractability A factor is worth listing when someone has a lever on it. That includes levers held by other teams and levers that are expensive — those get escalated or budgeted. It excludes facts of the universe: network partitions happen, hardware fails, a third party's region can go down. Those belong in the analysis as assumptions your design has to survive, and the *listed factor* is your system's response to them, not their existence. ## Grouping the survivors Sort what remains into the phases of the incident: 1. **Let the defect exist / reach production** — design, review, test, rollout policy. 2. **Delayed detection** — SLI coverage, thresholds, alert routing, noise. 3. **Delayed mitigation** — diagnosis difficulty, runbook, access, tooling speed. 4. **Amplified impact** — shared fate, missing bulkheads, client retry behaviour, cache stampedes. The grouping is diagnostic. An empty detection row means someone stopped after finding the bug. Four factors in "let the defect exist" and none in the other three phases means the analysis is about the code and not about the system, and the next different bug will produce an identically long outage. ## How many is right There is no magic number, but a well-pruned single-incident analysis usually lands around three to six factors with at least two phases represented. That is small enough to read and to fund, and large enough that it is not a chain wearing a list's clothing. ## In an interview Expect to be asked to prune out loud. Take two factors from the fourteen — one you keep, one you drop — and say exactly which filter each fails or passes. That is far more convincing than reciting the criteria.
- Isn't every contributing factor a counterfactual? Why is one form acceptable and the other not?Materiality is unavoidably counterfactual, but the two forms differ in what they can be checked against. A claim about system state — regional rollout would have capped impact at one region — is testable against the architecture and implies a concrete control. A claim about perception — the reviewer would have caught it — is only true because you now know the answer, cannot be falsified, and implies no control except more vigilance.
- How do you handle a factor that clearly mattered but sits entirely in another team's system?It stays in the analysis, because omitting it makes the document wrong. What changes is how it is carried: name it, quantify its contribution to impact, and route it to that team's own process rather than filing an item your team cannot execute. Alongside it, list what your service can do unilaterally to survive that factor next time — a timeout, a fallback, a bulkhead.
- What if pruning leaves only one factor?Then either the incident really was narrow — a single-user, quickly-detected, quickly-mitigated event — or, far more often, the analysis stopped early. I check the phase grouping first: one factor with nothing under detection or mitigation means nobody asked why it took as long as it did to notice and to fix, and those questions almost always yield something.
saying these in an interview costs you the question
- Lists every observation from the review as a contributing factor
- Keeps factors of the form "if only someone had noticed"
- Writes "insufficient testing" and calls it a finding
- Drops factors owned by another team because they are inconvenient
- Treats a long factor list as evidence of rigour