skip to content

How would you run a postmortem program across many teams so one team's outage produces learning and fixes beyond that team?

level: principalimportance: nice to knowfreq 33%

answer

  1. the theme, not the document
  2. comparable enough to query
  3. six local fixes = one platform gap
  4. curate, don't mandate reading
  5. never rank teams by incident count

basics

~20 s

Standardize a light template and store every postmortem in one searchable repository, then analyze across incidents for recurring themes, convert repeated local fixes into one platform-level fix with a funded owner, and circulate a small number of high-value write-ups rather than mandating that everyone read everything.

solid answer

~50 s

The unit of learning I care about at org scale is not the individual document — it is the *theme across documents*. So I would standardize enough of the template that incidents are comparable (impact, durations, contributing factors, typed action items) and store everything in one searchable repository with consistent tags. Then I would actually read across it quarterly: when six teams have each independently written "our config push had no staged rollout", that is not six team problems, it is one missing platform capability, and the right response is a funded platform fix rather than six local ones. Distribution matters too, but pushing every postmortem to everyone gets nothing read — a curated few high-value write-ups circulated widely works far better. And I would couple depth to severity, because mandating a full postmortem for every page produces ritual documents and burns the credibility of the whole practice.

go deeper

for a junior

Know that postmortems are stored centrally and searchable, and that checking whether a similar incident already happened is a legitimate first step during an incident, not a sign you should have known.

for a middle

Be able to say why a shared template and consistent tags matter: they make incidents comparable so questions can be asked across the whole corpus rather than one document at a time.

for a senior

Show that you look across incidents, not just at your own. Be ready to describe recognizing a theme in several teams' postmortems and arguing for one shared fix instead of repeating a local one.

for a principal

Own the whole program economics: what is mandated versus optional, who is funded to read across the corpus and act on themes, which metrics you publish, and why you refuse to rank teams by incident count.

## The scaling problem At one team, a postmortem is a document that changes that team's system. Across fifty teams, the same practice can produce five hundred documents a year and change nothing structurally, because each team fixes its own instance of a problem that is actually shared. The job of an org-level program is to convert local incidents into three things a single team cannot produce: **cross-incident themes**, **platform-level fixes**, and **transferred learning**. ## Standardize just enough Comparability is the whole reason to standardize. If impact is always quantified, durations always derivable from a timestamped timeline, and action items always typed, then you can ask questions across the corpus: which failure classes are most expensive, where is detection time worst, which services generate the most repeat incidents. If every team invents its own shape, the corpus is a pile of prose. The counter-pressure is real: a heavy mandated template gets filled in ritually. Mandate the small comparable core, let teams add whatever else helps them, and keep the document short enough to be written in a couple of hours. ## One searchable repository One home, indexed, with consistent tags: service, severity, failure class, whether it was a repeat. The high-value query is not "find the postmortem for last Tuesday" — it is "show me every incident in the last year where a config change caused customer impact". That query is what turns a document store into an analysis tool, and it is why per-team wikis and scattered documents fail at this scale. ## Read across, then fund a platform fix This is the part that actually pays. When the same contributing factor appears in six teams' postmortems, each team's local fix is six times the cost and covers a fraction of the surface. The org-level move is to name the theme, pick an owner with a real budget, and build the capability once — staged config rollout, a standard rollback path, a shared saturation-alerting library. The uncomfortable corollary: that platform work has to be funded from somewhere, and the argument for it lives in aggregated incident cost, not in any single postmortem. Which is why quantified impact in the template is non-negotiable — without it the theme analysis produces an interesting observation and no money. ## Distribute selectively Mandating that everyone read every postmortem produces compliance and no learning. What works is curation: a small number of genuinely instructive write-ups circulated widely — a "postmortem of the month"-style newsletter is the pattern Google's SRE book describes — plus reading groups for teams whose systems resemble the one that failed. The selection criterion is transferability, not severity: a mid-sized incident whose mechanism exists in twenty other services teaches more than the biggest outage of the year in a system nobody else runs. Also make the repository the first stop during an incident. "Has anyone seen this before?" answered by a search rather than by asking in chat is where the corpus repays its cost most directly. ## Program metrics A handful, and be careful with each: - **Time from incident to published postmortem**, by severity. - **Action-item completion rate and aging**, aggregated but *never compared between teams as a leaderboard* — the moment a number becomes a ranking, teams manage the number. - **Repeat-incident rate**: the share of incidents whose failure class already appears in the corpus. This is the closest thing to a direct measure of whether the program works. - **Postmortem count per team is not a quality metric.** A team with more incidents may be running riskier systems or simply being more honest about declaring them, and rewarding a low count buys you under-declaration. ## Couple depth to severity A sustainable program has tiers: full document plus review for customer-impacting incidents, a short structured record for minor ones, nothing for routine noise. Over-mandating is the most common way these programs die — engineers write documents to satisfy a process, quality collapses uniformly, and leadership concludes postmortems do not work. ## What it costs Be honest in an interview about the cost side: a program of this kind needs someone whose job includes reading the corpus, curating the newsletter and chasing themes; it needs funded platform capacity to act on what the themes reveal; and it needs leadership willing to not use the data punitively. Without the third, teams under-declare incidents and soften their documents, and every other part of the program degrades from the inside — which is a far more expensive failure than having no program at all.

  • What single metric best indicates whether the program is producing learning rather than paperwork?
    The repeat-incident rate — the share of incidents whose failure class already appears in the corpus. Publication counts and document quality scores measure activity; repeats measure whether prior learning actually changed the system. Expect it to be non-zero even in a healthy program, and watch its trend rather than its absolute value.
  • Why not rank teams by action-item completion rate to drive accountability?
    Because a ranked number gets managed. Teams commit to fewer items, close easy ones first, and declare fewer incidents — and you lose the honest data the program runs on. Aggregate the metric to spot systemic problems, and handle individual teams through conversation, where you can tell the difference between a team that is overloaded and one that is not trying.
  • How do you decide when a recurring theme deserves a platform fix rather than local ones?
    Compare aggregate cost against the cost of building it once. If several teams have each spent real incident time on the same missing capability and each local fix covers only its own service, the platform version is usually cheaper and more complete. The blocker is rarely the analysis — it is that the platform fix needs a funded owner, which is why quantified impact in every postmortem matters.
  • A team has far more postmortems than its peers. What do you conclude?
    Nothing on its own. It may run riskier or more complex systems, may sit at a failure-prone integration point, or may simply be more honest about declaring incidents. Treating the count as a quality signal reliably buys under-declaration, which costs more than the incidents did. Look at severity, repeat rate and impact instead, and treat the count as a prompt to ask questions.

saying these in an interview costs you the question

  • Mandates a full postmortem for every page that fires
  • Scatters postmortems across per-team wikis with no shared index
  • Ranks teams by incident count or completion rate
  • Requires everyone to read every postmortem
  • Identifies cross-team themes but funds no platform owner to fix them

context