skip to content

How soon after an incident should a postmortem be drafted and reviewed, and who should write it?

level: middleimportance: should knowfreq 44%

answer

  1. memory and retention both decay
  2. name the author during the incident
  3. days, not weeks, to draft
  4. responders write, an outsider reviews
  5. leave the meeting with owners and dates

basics

~20 s

Assign an author during the incident, draft within a few business days while memory and short-retention telemetry are still available, and review within a week or two. The responders write it; a reviewer outside the incident checks it is comprehensible and that the action items are real.

solid answer

~50 s

I assign the postmortem owner *during* the incident, usually the incident commander or a responder they name, so it does not become an orphan once the adrenaline is gone. The draft target is a few business days: responder recollection decays fast, chat scrollback and high-resolution metrics age out of short retention, and the longer the document waits the more the action items compete with newer work that feels more urgent. Authorship belongs to the people who responded — they are the only ones who know why a decision looked reasonable at the time — but the review should include at least one person who was not involved, because they are the ones who will catch jargon, missing context, and an action item that is really a wish. The review meeting should be short and it should end with owners and dates agreed out loud, not "we'll assign these later".

go deeper

for a junior

Know that postmortems are written within days, not weeks, and that the people who responded write them. If you are a responder, save timestamps, graph links and chat excerpts before they age out.

for a middle

Explain both decay curves — human recollection and telemetry retention — and why a late document quietly gets written from whatever logs survived. Say who should be in the review and what the meeting must produce.

for a senior

Demonstrate that you use publication speed as leverage: the days right after an outage are when a reliability item can win against roadmap work. Talk about assigning the author during the incident and coupling depth to severity.

for a principal

Own the standard: what draft and publish deadlines apply at each severity, who reviews across teams to keep quality even, and how you keep review meetings from degenerating into unread status readouts.

## Why the clock matters Everything a postmortem needs decays. Human recollection of *why* a decision looked correct at 03:00 is gone within about a week, and what remains is a tidied-up reconstruction that makes everyone look more rational than they were — which is exactly the information you least want to lose. Chat scrollback gets buried, dashboards roll off high-resolution retention into coarser rollups, and short-lived debug logs expire. A postmortem written three weeks later is frequently written *from the logs that survived*, which quietly biases the account toward whatever happened to be retained. The second decay is organizational. The impact of an outage is loudest in the days after it. That is when a P1 action item can win against roadmap work. Three weeks on, the outage is a story people tell, the team has new commitments, and the same item now has to displace something already promised. Speed of publication is, in practice, a lever on whether the fixes get funded at all. ## Assign the author during the incident The single most effective mechanic here is naming the postmortem owner while the incident is still open — the incident commander either takes it or names someone. An unassigned postmortem after an exhausting night is an orphan, and orphans are late. Naming it early also means the author starts collecting timestamps, screenshots and graph links *while they are still on screen*, which converts the write-up from an archaeology exercise into a transcription. A reasonable house standard: draft within three business days for a major incident, review within a week, published and closed within two. Numbers vary; having *a* number and tracking against it is what matters, because "soon" is not a deadline. ## Who writes it The responders. They have the only copy of the reasoning — what they believed, what signal they were looking at, what they ruled out and why. Handing the write-up to an uninvolved engineer produces a document that is accurate about events and empty about decisions, and it also removes the learning from the people best placed to use it. The common counter-argument is that responders will write self-servingly. In practice that risk is handled by the review, not by changing the author, and a document written by someone who was not there is not more objective — it is just less informed. ## Who reviews it A good review has three roles present: 1. **The responders**, who correct the timeline. 2. **A senior engineer familiar with the system**, who challenges whether the contributing factors are actually understood and whether the action items address them. 3. **At least one person with no involvement in the incident**, who is the real test of whether the document is comprehensible to the audience it was written for. If they cannot follow it, neither will the engineer who finds it in the repository in a year. Some organizations add a rotating reviewer or a small review group for higher-severity incidents specifically to keep quality consistent across teams. ## What the review meeting is for — and what it is not It is not a re-telling of the incident. Everyone reads the document first; if they have not, the meeting is worthless and should be rescheduled. The meeting exists to resolve disagreements about the timeline, to challenge weak action items, and — most importantly — to leave with every committed item carrying a named owner and a date. "We'll sort out owners after" is the sentence that produces an unowned backlog; the meeting is the moment when the room's attention makes assignment cheap, and that leverage is not recoverable later. A meeting that has become a twenty-person status readout has lost its function. Cut it to the people who can actually change something, or make the review asynchronous with a comment deadline and reserve the synchronous slot for the incidents where genuine disagreement exists. ## Not every incident needs the same weight Coupling postmortem depth to severity keeps the practice sustainable: a full document with a review meeting for the incidents that hurt customers, a short written record for the smaller ones, and nothing at all for routine noise. Demanding a full postmortem for every page produces ritual documents nobody reads and crowds out depth on the incidents that deserve it.

  • Why not have a neutral engineer who was not involved write the postmortem, for objectivity?
    Because objectivity is not the scarce resource — reasoning is. Only the responders can say what they believed at the time, what signal they were watching, and what they ruled out. An uninvolved author produces an accurate list of events and an empty account of decisions. Handle the self-serving risk in review, where an uninvolved reader challenges the document, rather than by removing the people who hold the information.
  • What should happen in the review meeting if the room disagrees about a contributing factor?
    Record the disagreement in the document rather than voting it away or letting the loudest person settle it. Then convert it into a time-boxed, owned item to resolve it. A postmortem that papers over an unresolved technical disagreement produces action items aimed at the wrong mechanism, and the disagreement resurfaces during the next incident at the worst possible moment.
  • Should every incident get a full postmortem with a review meeting?
    No — couple depth to severity. Full document and review for incidents with real customer impact, a short written record for minor ones, nothing for routine noise. Mandating full postmortems for every page produces documents written to satisfy the process, and the effort comes straight out of the incidents that actually deserved analysis.

saying these in an interview costs you the question

  • Says the postmortem can wait until the team has spare time
  • Leaves the postmortem unassigned when the incident closes
  • Hands the write-up to someone who was not involved, for "objectivity"
  • Ends the review meeting with owners to be assigned later
  • Runs the review as a re-telling for people who did not read the document

context