What sections does a standard incident postmortem document contain, and what is each section for?
answer
- impact before mechanism
- timestamped, not narrative
- detection and resolution stated separately
- went well / went wrong / got lucky
- action items table with owners
basics
~20 sA postmortem records customer impact and duration, a timestamped timeline, how the incident was detected and resolved, the contributing factors, lessons split into what went well / what went wrong / where we got lucky, and a table of owned, dated action items.
solid answer
~50 sMost house templates descend from the example postmortem in Google's SRE book, and they share the same spine. A **summary** and **impact** section state, up front, what customers experienced and for how long — that is the part executives read. A **timeline** of timestamped events gives you time-to-detect, time-to-mitigate and time-to-resolve for free. **Detection** and **resolution** say how you found out and what actually stopped the bleeding. An **analysis** section records the contributing factors. **Lessons learned** is deliberately split into what went well, what went wrong, and *where we got lucky* — the last one surfaces the coincidences that saved you and won't repeat. Finally, an **action items** table: one row per change, each with a type, a named owner, a priority and a tracker link. Everything above the action items exists to justify the rows below it.
code
markdown · 38 lines# Postmortem: checkout 5xx spike, 2026-02-14
Status: final | Authors: priya.n, tomas.k | Reviewed: 2026-02-18
## Summary
A config push disabled connection pooling on the checkout client; checkout returned 5xx for 41 minutes.
## Impact
41 minutes of elevated errors, ~12,000 failed checkouts, 3 enterprise tickets.
## Timeline (UTC)
- 09:14 config push rolled out to all regions
- 09:19 error-rate SLO burn alert fires
- 09:22 on-call acknowledges
- 09:31 incident declared, sev2
- 09:47 rollback started
- 09:55 error rate normal
## Detection
SLO burn-rate alert. No customer report preceded it.
## Resolution
Rollback of the config push. No code change was required.
## Contributing factors
- The pooling setting had no staged rollout.
- Connection saturation was not a paging signal.
## Lessons learned
**Went well:** rollback path was rehearsed and took 8 minutes.
**Went wrong:** 17 minutes from alert to declaration.
**Got lucky:** the push landed at off-peak traffic; at peak this would have been sev1.
## Action items
| Action | Type | Owner | Priority | Tracker |
|---|---|---|---|---|
| Stage config pushes region by region | prevent | tomas.k | P1 | REL-482 |
| Alert on connection-pool saturation | detect | priya.n | P1 | REL-483 |go deeper
Be able to name the sections and say what each is for, especially impact, timeline and action items. Practice writing an impact line with a real number and a duration in it rather than "the service was degraded".
Explain why the timeline is timestamped: it turns detect, mitigate and resolve into measurable durations, and the biggest gap tells you which kind of action item to write. Be ready to defend the three-way lessons split.
Show that you use the template to make incidents comparable across time — searchable impact figures and consistent durations are what let you argue for reliability work later. Talk about what you deliberately cut from a bloated template.
Own the tradeoff between a template rich enough to support cross-incident analysis and short enough that engineers fill it in honestly. Be ready to say which sections you would make mandatory org-wide and which you would leave to teams.
## What the document is actually for A postmortem is not a report filed to prove that an incident was taken seriously. It is a mechanism with two outputs: a shared, accurate account of what happened that anyone in the company can read months later, and a short list of changes that make the next occurrence less likely, less severe, or faster to catch. Every section of the template earns its place by feeding one of those two outputs. If a section feeds neither, cut it — long templates get filled in ritually and read by nobody. The canonical section list most companies use descends from the example postmortem published in Google's SRE book (2016). Names vary between houses; the spine does not. ## Metadata and summary Title, date, authors, status (draft / in review / final), and a two-or-three-sentence summary. The status field matters more than it looks: a postmortem that sits in "draft" forever is one of the classic failure signatures, and having the state on the document makes that visible. ## Impact What customers experienced, quantified, and for how long. "Checkout returned 5xx for 41 minutes; approximately 12,000 requests failed; 3 enterprise customers opened tickets." This section comes first because it is what non-engineers read, and because it is the number that later justifies — or fails to justify — the cost of the action items. An impact section written as "the service was degraded" is not usable for prioritization. ## Timeline A list of timestamped events, in UTC or one clearly stated zone: first error, first alert fired, human acknowledged, incident declared, mitigation applied, recovery confirmed, incident closed. The timeline is the raw material for the durations you care about — time to detect, time to mitigate, time to resolve — and those durations are usually where the improvable gap is. A team that took eight minutes to fix a problem it took ninety minutes to notice has a detection problem, not a fix problem, and only a timestamped timeline makes that obvious. Narrative entries like "we recovered sometime in the afternoon" destroy the section's value. ## Detection and resolution How you learned about it (an alert, a customer ticket, an engineer noticing a graph) and what actually restored service. Separating these from the timeline forces an honest answer to "would we have caught this without the customer telling us?" ## Analysis: contributing factors What combined to produce the outage. Keep it factual and mechanism-level. How you *conduct* that analysis — the questioning technique, trigger versus underlying cause, why one "root" cause is usually a fiction — is a discipline of its own and is not what the template section is for; the section is just where its output lands. ## Lessons learned: went well / went wrong / got lucky The three-way split is not decoration. **What went well** protects the practices that worked so a later reorganization does not delete them by accident. **What went wrong** is the source list for action items. **Where we got lucky** is the section people skip and the one that most often predicts the next outage: the failure landed at 03:00 when traffic was one-tenth of peak; the one engineer who knew the subsystem happened to be awake; the corrupted replica happened to be the one that was already out of rotation. Luck is not a control. Anything recorded here is a risk you are still carrying, and it frequently converts directly into an action item. ## Action items A table, not a paragraph. Each row: the change, a type (commonly *prevent*, *mitigate*, *detect*, or *process*), a named individual owner, a priority, and a link into the tracker where the work actually lives. The table is the only part of the document that changes the system; everything above it is the evidence that the rows are worth doing. ```markdown | Action | Type | Owner | Priority | Tracker | |---|---|---|---|---| | Add jittered retry budget to checkout client | prevent | priya.n | P1 | REL-482 | | Alert on replica lag > 30s | detect | tomas.k | P1 | REL-483 | | Add "drain region" step to the runbook | mitigate | priya.n | P2 | REL-490 | ``` ## Supporting information Graphs, log excerpts, chat transcript links. Useful, and deliberately last — the document must be readable without it. ## What a template cannot do A template standardizes the *shape* of the record so it is searchable and comparable across dozens of incidents. It does not make the content honest, does not make the analysis deep, and does not make the action items happen. Teams that adopt a template and see no change in reliability have usually adopted only the shape.
- Why is "where we got lucky" worth a section of its own rather than a line in what went well?Because they are opposite signals. "Went well" records controls you built and should protect; "got lucky" records risks you are still carrying that happened not to fire this time — off-peak timing, one expert being awake, a replica that was already out of rotation. Merging them lets an unmitigated risk masquerade as a working control, which is exactly the confusion that produces a repeat outage at a worse hour.
- What do you get from the timeline that you cannot get from the narrative sections?Durations. Timestamps give you time-to-detect, time-to-acknowledge, time-to-mitigate and time-to-resolve as arithmetic rather than opinion, and the largest gap tells you where to spend. A team that mitigated in eight minutes but noticed after ninety has a detection problem; without timestamps that conclusion is a matter of debate, and the action items usually end up aimed at the wrong stage.
- Should the impact section be written in engineering terms or customer terms?Customer terms first, engineering terms second. "Checkout failed for 41 minutes, roughly 12,000 orders affected" is what non-engineers read and what later justifies the cost of the action items. Error rates and saturation graphs belong in the timeline and supporting information. An impact section that says only "the service was degraded" cannot be used to prioritize anything against feature work.
Think of it as an accident report plus a work order: the narrative sections are the evidence, and the action-item table is the only part that changes anything.
saying these in an interview costs you the question
- Says the postmortem is a report for management, not a change mechanism
- Writes a narrative timeline with no timestamps
- States impact as "the service was degraded" with no duration or count
- Treats "where we got lucky" as the same thing as what went well
- Leaves action items as prose in the document instead of a table with owners