skip to content

Templates, Action Items & Follow-Through

A postmortem is only as good as the changes it produces. Interviewers ask how you make action items actually happen, because unowned, untracked follow-ups are the most common postmortem failure mode.

on this pageshow

questions

5

What separates a good postmortem action item from a bad one?

level: middleimportance: must knowfreq 72%

answer

  1. can an outsider tell it is done?
  2. investigations have no definition of done
  3. a person, not a team alias
  4. lives in the backlog, not the doc
  5. prevent, mitigate, detect — not all prevent

basics

~20 s

A good action item names one concrete change, has a single named human owner, a priority and a due date, lives in the team's normal tracker, and has an unambiguous definition of done. Bad items are open-ended investigations or exhortations to be careful.

solid answer

~50 s

The test I apply is: could someone other than the author tell, six weeks from now, whether this is finished? That rules out three common shapes. "Investigate why the cache stampeded" has no completion criterion — an investigation is a task you do *during* the postmortem, not an outcome of it. "Be more careful when editing production config" asks humans to change without changing the system, so it will not survive the next tired on-call. "Improve monitoring of the payments service" names no specific signal. A good item names the change, an individual owner rather than a team alias, a priority, a date, and a tracker ID so it competes for capacity alongside everything else. I also type items — commonly prevent, mitigate, detect, process — because a list that is entirely *prevent* leaves you no faster at catching or containing the next failure, which will be a different one.

code

yaml · 27 lines
yaml
# Weak items: nothing here can be marked done six weeks from now.
bad:
  - title: Investigate why the connection pool saturated
  - title: Be more careful when editing production config
  - title: Improve monitoring of the payments service
    owner: platform-team

# Strong items: one change each, an individual owner, a date, a tracker row.
good:
  - title: Stage config pushes region by region
    type: prevent
    owner: tomas.k
    priority: P1
    due: 2026-03-06
    tracker: REL-482
  - title: Alert when replica lag exceeds 30s for 5 minutes
    type: detect
    owner: priya.n
    priority: P1
    due: 2026-03-06
    tracker: REL-483
  - title: Rehearse the region-drain runbook end to end once
    type: mitigate
    owner: priya.n
    priority: P2
    due: 2026-03-20
    tracker: REL-490

go deeper

for a junior

Know the fields every action item needs: a concrete change, a named person, a priority, a date, and a ticket. Be able to spot "be more careful" as a non-item and say what the systemic version would be.

for a middle

Explain the definition-of-done test and why an open-ended investigation cannot pass it. Be ready to rewrite a vague item into a buildable one on the spot, naming the specific signal or the specific guardrail.

for a senior

Show judgment about the mix: argue when a mitigation that halves recovery time beats a prevention aimed at the cause you just saw, and be honest about capping the list at what you will actually staff.

for a principal

Own the accounting. Be ready to say how action-item capacity is reserved against roadmap pressure, what your standard is for closing an item, and why an honest "considered and not taken" beats an open ticket nobody intends to do.

## The single test Every quality rule for action items collapses into one question: **could a person who was not in the incident tell, weeks later, whether this item is done?** If the answer is no, the item will quietly rot in the document, and the postmortem will have produced a nice narrative and no change. ## The three shapes that always fail **The open-ended investigation.** "Investigate why the connection pool saturated." Investigation has no definition of done — you can always investigate more — so the item can never be closed honestly, only abandoned. Worse, an unfinished investigation is usually a sign the analysis was cut short and the postmortem was published before it was understood. If you genuinely need more digging, the item should be time-boxed and produce an artifact: "Spend up to two days reproducing the saturation and either file the fix ticket or write up why it is not reproducible, by 6 March." Now it can close. **The exhortation.** "Be more careful when editing production configuration." "Remember to check the dashboard after a deploy." These ask a human to be reliably different next time, under worse conditions than the ones in which they already failed. They are also, in practice, a fault attribution wearing procedural clothes. The systemic version of the same item is what you want: make the dangerous edit impossible, require staged rollout, add a confirmation for the destructive path, or add the check to an automated gate. **The unbounded improvement.** "Improve monitoring of the payments service." There is no signal named, no threshold, no owner of the judgment. Compare: "Alert when replica lag exceeds 30 seconds for 5 minutes, routed to the payments rotation." That one is buildable and checkable. ## The fields a real item carries - **A concrete change**, phrased as an outcome, one item per row. Two changes bundled into one row means half of it silently never happens. - **A single named individual owner.** Not "the platform team", not "SRE". A team alias has no calendar and no accountability; assigning to a team is how items become nobody's. The individual can delegate, but someone answers for it. If no individual will take it, that is real information: the item is not actually going to happen and you should either find an owner or drop it honestly. - **A priority**, on the same scale the team uses for everything else, so it can be compared with feature work rather than living on a separate, ignorable list. - **A due date**, which gives you an aging signal. An item three weeks past its date is a visible fact; an undated item is never late. - **A tracker ID.** The item's canonical home is the team's backlog. The postmortem *links* to it. Action items that live only inside the document never enter sprint planning, and anything that does not enter planning does not get staffed. - **A type.** Google's SRE book template classifies items as prevent, mitigate, detect or process. ```yaml action_items: - title: Stage config pushes region by region type: prevent # prevent | mitigate | detect | process owner: tomas.k # an individual, never a team alias priority: P1 due: 2026-03-06 tracker: REL-482 # lives in the backlog; the doc only links here ``` ## Why the typing matters A set of action items that is entirely *prevent* is a common and expensive mistake. Prevention items are aimed at the failure you just had — but the next incident will be a different one, and the prevention you just built will not apply to it. *Detect* and *mitigate* items generalize: a faster rollback path, a saturation alert, a documented drain procedure all pay out across failure modes you have not seen yet. A healthy set usually has at least one of each, and if you can only afford one item, a mitigation that halves your recovery time is often worth more than a prevention that eliminates one specific cause. *Process* items — updating a runbook, changing an escalation path — are the cheapest and the easiest to fake. A runbook step that has never been executed is not a control, so a process item ideally ends with someone having actually run it. ## Volume discipline Twenty action items is not a thorough postmortem; it is an unstaffed wish list, and shipping none of twenty is the normal outcome. Commit to the handful you will genuinely staff, mark the rest explicitly as "considered and not taken" with a one-line reason, and let that honest record be the thing a future reader finds — rather than a graveyard of open tickets that teaches everyone the items are decorative.

  • Why insist on an individual owner rather than assigning an item to the owning team?
    Teams have no calendar and no accountability. An item assigned to "platform" is nobody's on the day sprint work is chosen, and it produces no aging signal that anyone feels. A named owner can still delegate, but someone answers for it in review. The failure to find any individual willing to own an item is itself useful information — it usually means the item is not really going to be done.
  • When is an "investigate further" action item legitimate?
    When it is time-boxed and produces a named artifact. "Spend up to two days reproducing the saturation, then either file the fix ticket or write up why it isn't reproducible, by 6 March" has a definition of done. An open-ended "investigate" does not, and usually signals the postmortem was published before the incident was understood.
  • You can fund exactly one action item this quarter: eliminating the specific cause, or halving rollback time. Which?
    Usually the rollback. Prevention pays out only against the failure you already had; the next incident is a different one. Halving recovery time reduces the impact of every future failure mode, including the ones you have not imagined. I would flip that choice if the specific cause is likely to recur soon and its blast radius is severe enough that recovery speed does not save you.
  • How many action items should a postmortem produce?
    As many as you will genuinely staff — typically a handful. A list of twenty is an unstaffed wish list, and closing none of twenty teaches the team the items are decorative. Record the rejected candidates explicitly as considered-and-not-taken with a one-line reason, so a future reader sees a decision rather than a graveyard of stale tickets.

saying these in an interview costs you the question

  • Files "investigate X further" as a completed-postmortem action item
  • Assigns items to a team alias instead of a named person
  • Writes "be more careful" instead of changing the system
  • Leaves action items only in the document, never in the backlog
  • Produces twenty items and staffs none of them
  • Makes every item a prevention for the exact failure just seen

context

open as a page

What sections does a standard incident postmortem document contain, and what is each section for?

level: juniorimportance: should knowfreq 68%

basics

~20 s

A postmortem records customer impact and duration, a timestamped timeline, how the incident was detected and resolved, the contributing factors, lessons split into what went well / what went wrong / where we got lucky, and a table of owned, dated action items.

open as a page

How soon after an incident should a postmortem be drafted and reviewed, and who should write it?

level: middleimportance: should knowfreq 44%

basics

~20 s

Assign an author during the incident, draft within a few business days while memory and short-retention telemetry are still available, and review within a week or two. The responders write it; a reviewer outside the incident checks it is comprehensible and that the action items are real.

open as a page

Six months after an outage, the same class of failure recurs and you discover the earlier postmortem's action items were never completed. How do you fix follow-through?

level: seniorimportance: should knowfreq 56%

basics

~20 s

Move action items out of the document and into the team's normal backlog with individual owners, priorities and dates, then measure completion rate and aging as a standing metric, review overdue items regularly, and use the error-budget policy to buy the capacity when reliability work keeps losing to roadmap work.

open as a page

How would you run a postmortem program across many teams so one team's outage produces learning and fixes beyond that team?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Standardize a light template and store every postmortem in one searchable repository, then analyze across incidents for recurring themes, convert repeated local fixes into one platform-level fix with a funded owner, and circulate a small number of high-value write-ups rather than mandating that everyone read everything.

open as a page