skip to content

Six months after an outage, the same class of failure recurs and you discover the earlier postmortem's action items were never completed. How do you fix follow-through?

level: seniorimportance: should knowfreq 56%

answer

  1. four different failures hide behind "incomplete"
  2. the backlog, not the document
  3. aging beats completion rate
  4. reserved capacity, defended
  5. budget policy pre-wins the argument

basics

~20 s

Move action items out of the document and into the team's normal backlog with individual owners, priorities and dates, then measure completion rate and aging as a standing metric, review overdue items regularly, and use the error-budget policy to buy the capacity when reliability work keeps losing to roadmap work.

solid answer

~60 s

First I would separate the two possible failures, because the fixes differ. If the items were *never staffed*, the problem is that reliability work has no protected capacity, and the answer is structural: items land in the same backlog as feature work with individual owners and dates, a fixed share of each sprint is reserved for them, and if that keeps getting eaten, the error-budget policy is the lever that converts a reliability deficit into an actual claim on engineering time. If the items were *staffed but wrong* — closed, yet the failure recurred — the problem is item quality or analysis depth, not tracking. Then I would make the state visible: track completion rate and the age of open P1 items from past incidents as a team metric, review overdue items at the same cadence you review anything else, and cap what you commit to so the number means something. And I would say plainly in the new postmortem that a prior one predicted this — that fact is the strongest argument you will ever have for funding the fix.

go deeper

for a junior

Know that action items belong in the normal ticket tracker, with an owner and a date, and that the postmortem only links to them. If you own one, say early when you cannot get to it rather than letting it age silently.

for a middle

Explain the mechanics you would put in place: tagging items so they are queryable as a set, tracking aging as well as completion, and reviewing overdue items on a fixed cadence with someone who can reprioritize.

for a senior

Distinguish the never-staffed case from the completed-but-ineffective case before proposing a fix, and show that you use reserved capacity and the error-budget policy as levers rather than relying on people trying harder.

for a principal

Own the funding question: how much engineering capacity is structurally reserved for reliability work, what pre-agreed policy redirects it when a service is unreliable, and how you make a recurrence visible to the people who set priorities.

## Diagnose before you prescribe "The action items were never completed" hides at least four different failures, and the interviewer is usually checking whether you distinguish them: 1. **They never left the document.** The items were written in the postmortem and nowhere else, so they never appeared in any planning conversation. This is the most common case and the easiest to fix. 2. **They were tracked but never staffed.** Tickets exist, aged, and lost every prioritization round to work with a nearer deadline. 3. **They were completed but ineffective.** Closed on time, and the failure recurred anyway — which means the analysis was wrong or the items addressed a symptom. 4. **They were the wrong items to begin with** — twenty of them, none funded, so "incomplete" was the predictable outcome of over-committing. Only the first two are follow-through problems. Cases three and four are quality problems that better tracking will not touch. ## Make the backlog the single home Action items must live where work is chosen. The postmortem *links* to tracker rows; it does not store them. Anything held only in a document is invisible during sprint planning, and invisible work is unstaffed work. The corollary is that they carry the same fields as everything else — owner, priority, due date — so they can be compared with feature work rather than sitting on a separate list that is easy to ignore. Tag them so they are recoverable as a set. A saved query for "open, priority P1, source = postmortem, older than 30 days" is the entire mechanism behind everything else in this answer. ## Measure two numbers, not one - **Completion rate** of committed items over a rolling window. - **Age of open high-priority items** — the more diagnostic of the two. A team can show a healthy completion rate while its hardest, most valuable items sit untouched for months, because the easy ones close and flatter the average. Both are only meaningful if you cap what you commit to. If a postmortem produces twenty items and you staff four, the honest record says four were committed and sixteen were considered and declined with a reason — not that you completed 20% of your commitments. ## Reserve the capacity, then defend it The real reason items rot is that they compete with work that has a customer waiting. Three mechanisms help, in increasing order of strength: - **A standing share of each sprint** reserved for reliability work. Cheap to agree, easy to erode. - **A recurring review of overdue items** with the people who can reprioritize in the room. Aging items that are read out loud regularly either get done or get closed honestly; both beat silence. - **An error-budget policy.** This is the strongest lever, because it is a pre-agreed rule rather than a negotiation held while under pressure: when a service has burned its budget, reliability work — including outstanding postmortem items — takes precedence. The value is that the argument was won *before* the moment when nobody wants to have it. ## Close the loop honestly An item that is not going to be done should be closed with a stated reason, not left open. A backlog of permanently open postmortem items teaches everyone that the items are decorative, and that lesson generalizes — the next postmortem's items get written with less care because the authors already know how this ends. ## Use the recurrence itself When the same class of failure returns, the new postmortem should state it explicitly: an earlier postmortem identified this risk, item REL-482 was opened and never staffed, and here is the cost that decision has now produced. This is not blame-hunting — it is the single most persuasive artifact you will ever have for funding reliability work, because it converts a hypothetical into a bill that has already been paid twice. Follow it with the prevention question one level up: what would have made that item survive prioritization? Usually the answer is that it needed protected capacity or a named owner with the authority to schedule it, and that is the fix worth writing down. ## What good looks like A team with healthy follow-through can answer, from a saved query and without preparation: how many action items are open, how old the oldest P1 is, and which ones were consciously declined. If nobody can answer that in a minute, the process is not running regardless of how good the documents look.

  • Why is the age of open items more diagnostic than the completion rate?
    Because completion rate is flattered by easy items. A team can close every runbook tweak on time, keep a 75% rate, and leave the two structural fixes untouched for eight months — and those two are the ones that would have prevented the recurrence. The age of the oldest open P1 exposes exactly what the average hides, and it is the number I would put on a dashboard.
  • Is it acceptable to close a postmortem action item without doing it?
    Yes, if you close it with a stated reason and a decision-maker. An honest "not doing this, the risk is accepted because X" is a real record a future reader can evaluate. What corrodes the practice is items left open indefinitely: they teach the team that action items are decorative, and the next postmortem gets written with correspondingly less care.
  • The team completed every action item and the failure still recurred. What does that tell you?
    That this is not a follow-through problem at all — it is an analysis or item-quality problem. The items addressed something adjacent to the real mechanism, or they were process changes that never became controls. The response is to reopen the analysis with people who were not in the first one, not to add more tracking around a process that already worked.
  • How do you raise a never-staffed action item in the new postmortem without it reading as blame?
    State it as a decision with a cost, not as a person's failure: the risk was identified, item REL-482 was opened, it lost prioritization for two quarters, and the resulting outage cost this much. Then ask the systemic question — what would have let that item survive prioritization? That framing keeps the document candid and turns the recurrence into the strongest funding argument you will ever have.

saying these in an interview costs you the question

  • Treats every incomplete item as a discipline problem rather than a capacity one
  • Keeps action items only in the postmortem document
  • Reports completion rate while the oldest P1 has aged for months
  • Leaves undoable items open forever instead of closing them with a reason
  • Commits to twenty items and staffs four, then calls it 20% completion

context