skip to content

Postmortems & Learning

Turning incidents into systemic improvement through blameless postmortems, contributing-factor analysis, and tracked action items. Interviewers use postmortem questions to gauge engineering culture and whether you actually learn from failure.

on this pageshow

questions

16

In incident analysis, what is the difference between the trigger of an outage and its underlying cause, and why does the distinction change what you fix?

level: juniorimportance: must knowfreq 50%

answer

  1. what started it vs why it was enough
  2. latent condition, present all along
  3. the last hundred deploys were fine too
  4. cheap gate now, structural fix later
  5. triggers set exposure frequency

basics

~20 s

The trigger is the event that started this outage; the underlying cause is the latent condition that made the trigger dangerous. Removing the trigger prevents this instance, fixing the latent condition prevents the whole class — and they cost very different amounts.

solid answer

~50 s

The trigger is what happened at t0 — a deploy, a config push, a traffic spike, a node dying. The underlying cause is the condition that had been sitting in the system harmlessly and turned that ordinary event into an outage: an unbounded queue, a retry path with no backoff, a leak that only matters after fourteen days of uptime. The distinction matters because it maps to two different fixes with two different price tags. Banning or gating the trigger is usually cheap and fast — freeze that config path, add a validation check, throttle the caller. Fixing the latent condition is usually expensive and slow, but it is the only thing that stops the next trigger you haven't thought of. Good practice is to do the cheap one now to stop the bleeding and be explicit that you are accepting the latent condition until the real fix lands, rather than pretending the incident is closed because the trigger is gone.

go deeper

for a junior

Be able to state plainly that the trigger is the event that started the outage and the underlying cause is the pre-existing condition that made that event dangerous, and give one concrete pair.

for a middle

Explain why the latent condition survived the previous hundred harmless deploys, and describe what changes when you fix the trigger versus the condition.

for a senior

Show the allocation decision: which fix you ship in the first week, which structural item you fund, and how you make carrying the latent condition an explicit, owned risk rather than an oversight.

for a principal

Own the standard that says a postmortem closing with only a trigger fix is incomplete, and hold the portfolio view of knowingly-carried latent conditions across services so they are decided on rather than forgotten.

## Two different questions "What started it?" and "why was that enough to break us?" are separate questions, and postmortems that answer only the first are the most common weak postmortem in circulation. The **trigger** is the proximate event: a release went out, a certificate expired, a partner sent ten times its usual volume, a hypervisor rebooted. Triggers are usually specific, well-timestamped, and easy to identify — they are what shows up first when you overlay the change log on the SLI graph. The **underlying cause** (also called the latent condition or contributing condition) is the property of the system that had to be true for that trigger to produce user impact. It was there before the incident, it was there during the last hundred deploys that went fine, and it is still there after you revert. ## A worked example A routine config push takes the service down. Overlay the graph: the push is the trigger, timestamps line up, done. Except: - the service parses config at startup and **exits** on a parse failure, rather than logging and keeping the last known-good config - the config is delivered to **all regions simultaneously**, so every replica hit the same bad file within seconds - the health check reports the process as healthy for 30 seconds after start, so the orchestrator kept replacing pods into the same failure The bad config is the trigger. Fail-closed parsing, global simultaneous delivery, and a health check that lies are the underlying causes. Ban the specific bad field and you are safe from that exact push. Any of the three underlying conditions will convert the next imperfect config into the same outage. ## Why triggers still matter The usual overcorrection is to declare triggers uninteresting. They are not, for two reasons. First, **trigger frequency is a real dial**. If the same class of trigger fires weekly and the latent fix is a quarter of engineering time, reducing trigger frequency — validation in CI, a staged config rollout, rate-limiting the noisy caller — buys you most of the risk reduction immediately and cheaply. Second, some triggers are themselves the finding. A certificate expiring is a trigger, but "nothing tracks certificate expiry" is a systemic gap wearing a trigger's clothes. ## The decision the distinction forces After the incident you have to allocate finite engineering time, and the choice is genuinely contested: | Option | Cost | Buys you | |---|---|---| | Remove or gate the trigger | Hours to days | This exact incident cannot repeat | | Reduce trigger frequency | Days | Fewer exposures to the same latent condition | | Fix the latent condition | Weeks | The whole class of triggers becomes survivable | | Reduce blast radius | Days to weeks | Any trigger hurts a fraction of users, not all | The honest answer in most postmortems is a fast trigger fix now plus one structural item, with a named owner and a date, and an explicit statement that the latent condition is being knowingly carried in the meantime. A postmortem that lists only the trigger fix and closes is how the same outage arrives six weeks later with a different first line. ## Naming it precisely Be careful with the phrase "root cause" in this context. Many teams use it to mean the latent condition, some use it to mean the trigger, and the ambiguity produces arguments that are really about vocabulary. Saying "the trigger was X; the conditions that made X an outage were Y and Z" is unambiguous, and it signals in an interview that you have thought about this beyond the label. ## In an interview Use a concrete pair from your own experience. "The trigger was a partner doubling their request rate on a Monday; the underlying cause was that our retry policy had no jitter, so their retries synchronized and we amplified their spike into a thundering herd against the database." Then say what you fixed first and why — that is the part that shows judgment rather than vocabulary.

  • The same latent condition has been in production for two years without incident. Does that make it low risk?
    It makes its trigger rate low so far, which is not the same thing. Risk is exposure times consequence, and the consequence side is unchanged — when it does fire it takes the service down. What actually raises the risk is anything that increases trigger frequency: a new caller, higher traffic, a change in deploy cadence. Two quiet years is evidence about the past environment, not about the next one.
  • How do you decide between fixing the latent condition and just reducing the blast radius?
    By cost and by how many trigger classes each covers. Fixing the condition removes one specific failure mode; reducing blast radius — regional staging, cells, a fail-static fallback — caps the damage from failure modes you have not thought of yet, which is usually the better bet when the latent fix is a multi-week rewrite. In practice I ship the blast-radius change first and schedule the deeper fix.

saying these in an interview costs you the question

  • Calls the trigger the root cause and closes the postmortem
  • Assumes reverting the change means the risk is gone
  • Treats triggers as uninteresting once a latent cause is found
  • Believes a latent condition that never fired is not a real defect
  • Uses "root cause" without saying which of the two is meant

context

open as a page

What separates a good postmortem action item from a bad one?

level: middleimportance: must knowfreq 72%

basics

~20 s

A good action item names one concrete change, has a single named human owner, a priority and a due date, lives in the team's normal tracker, and has an unambiguous definition of done. Bad items are open-ended investigations or exhortations to be careful.

open as a page

During a postmortem your 5 Whys chain ends at "the config parser didn't validate the field." Why do experienced reviewers push back on stopping there, and what analysis do you do instead?

level: middleimportance: must knowfreq 65%

basics

~20 s

5 Whys follows one causal line and stops at the first fixable defect, so it names the bug but not the review, test, rollout and detection gaps that let the bug reach production. Ask "what else" at each step and record multiple contributing factors.

open as a page

In a blameless postmortem, what does "blameless" actually mean, and how does a team still hold anyone accountable?

level: middleimportance: must knowfreq 76%

basics

~20 s

Blameless means the postmortem asks how the system let a reasonable action cause harm, not who to punish. Accountability survives as named owners for the fixes and an obligation to explain honestly; deliberate recklessness is a management matter handled outside the document.

open as a page

A draft postmortem lists the root cause as "human error — the operator ran the wrong command." Why would an SRE reject that, and what should the document say instead?

level: middleimportance: must knowfreq 62%

basics

~20 s

"Human error" is where an investigation stopped, not an explanation. Treat it as a symptom and record what made the wrong command reachable, plausible and undetected — no confirmation prompt, identical staging and production shells, a stale runbook, time pressure — and fix those.

open as a page

What sections does a standard incident postmortem document contain, and what is each section for?

level: juniorimportance: should knowfreq 68%

basics

~20 s

A postmortem records customer impact and duration, a timestamped timeline, how the incident was detected and resolved, the contributing factors, lessons split into what went well / what went wrong / where we got lucky, and a table of owned, dated action items.

open as a page

How soon after an incident should a postmortem be drafted and reviewed, and who should write it?

level: middleimportance: should knowfreq 44%

basics

~20 s

Assign an author during the incident, draft within a few business days while memory and short-retention telemetry are still available, and review within a week or two. The responders write it; a reviewer outside the incident checks it is comprehensible and that the action items are real.

open as a page

How do you reconstruct the timeline of an incident for its postmortem, and which timestamps in that timeline matter most?

level: middleimportance: should knowfreq 60%

basics

~20 s

Rebuild the timeline from machine evidence — deploy records, alert history, chat transcripts, dashboards — normalized to UTC. The load-bearing timestamps are impact start, detection, human engagement, first mitigation and impact end, because their gaps measure detection and mitigation separately.

open as a page

Six months after an outage, the same class of failure recurs and you discover the earlier postmortem's action items were never completed. How do you fix follow-through?

level: seniorimportance: should knowfreq 56%

basics

~20 s

Move action items out of the document and into the team's normal backlog with individual owners, priorities and dates, then measure completion rate and aging as a standing metric, review overdue items regularly, and use the error-budget policy to buy the capacity when reliability work keeps losing to roadmap work.

open as a page

A draft postmortem lists fourteen contributing factors for one outage. How do you decide which ones belong in the final analysis and which are noise?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Keep a factor if changing it would plausibly have prevented or materially reduced the impact, it can recur beyond this incident, and someone has a lever on it. Drop factors visible only in hindsight and factors so generic they would fit any postmortem.

open as a page

How would you tell whether a team's postmortem process is genuinely blameless rather than blameless only on paper?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Watch behaviour under cost, not policy. Do people volunteer that they were the one who acted, write their own postmortems, and declare incidents early? Do remediations target systems rather than individuals? Blame that has moved out of the document and into private conversations is the usual tell.

open as a page

For six months your team has skipped a required pre-deploy verification step because it is slow, and nothing has broken. What is this pattern called, and what do you do about it?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Normalization of deviance: a shortcut that keeps working quietly becomes the standard, and the margin it consumed stays invisible until it runs out. Close the gap between written and actual practice — make the step cheap enough to keep, or formally retire it after an explicit risk decision.

open as a page

Three outages this quarter had different proximate causes, but all three involved the same shared configuration service. How do you approach the analysis across them?

level: principalimportance: should knowfreq 30%

basics

~20 s

Analyze the set, not the incidents. Tag contributing factors across postmortems, find the factor that repeats, normalize for how much traffic that component actually carries, then argue a structural investment using aggregate impact rather than three separate small fixes.

open as a page

A canary deploy caught a fatal bug before any customer saw an error. Do you analyze it like an incident, and what would that analysis look at?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Yes, when the margin was thin or the save was luck rather than design. The analysis asks how close it came, whether the control that caught it was intentional and repeatable, and what would have happened had it fired ten minutes later.

open as a page

How would you run a postmortem program across many teams so one team's outage produces learning and fixes beyond that team?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Standardize a light template and store every postmortem in one searchable repository, then analyze across incidents for recurring themes, convert repeated local fixes into one platform-level fix with a funded owner, and circulate a small number of high-value write-ups rather than mandating that everyone read everything.

open as a page

After a customer-visible outage, an executive asks who was responsible and wants consequences. How do you respond as the engineering lead?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Give the executive what they actually want — confidence it will not recur — rather than a name. Present the missing controls, the fixes with owners and dates, and the cost of punishing: you lose your best-informed engineer and teach everyone else to hide the next one. Reserve genuine recklessness for a private management track.

open as a page