Walk me through a postmortem you wrote after an incident you caused.
answer
- the incident, ownership in one line
- blameless but not authorless
- root cause versus contributing factors
- action items with named owners
- what landed, what changed beyond you
basics
~20 sProbes whether you can analyse your own mistake without either flinching or self-flagellating. Answer with a real write-up: root cause held separate from contributing factors, actions with named owners, and an honest count of what actually shipped.
how to answer
5 beats- the incident in two sentences, ownership statedCompress the outage itself hard, since the question is about the document. Say what broke, roughly how bad it was, and that the change was yours, then move on within about fifteen percent of your airtime.
- how you kept it blameless while still owning itExplain the register you chose: naming exactly what the change did while asking why the system permitted it. Mention that being the author of the change makes it easier to set that tone for everyone else.
- root cause held apart from contributing factorsThis is the analytical core and deserves the largest share of your airtime. State the single mechanism, then list the conditions separately: missing ownership, stale documentation, a lagging alert. Do not let the list absorb the cause.
- the actions, their owners, and what really landedGive the count, who owned each item, and the honest completion picture including anything you renegotiated or dropped. One or two sentences on the item that slipped buys more credibility than a clean claim.
- what the document changed for other peopleClose on effect beyond the fix: a review that catches this class now, a norm others adopted, a number showing the difference. Keep this to the last fifth of your airtime and make it specific.
your answer
4 story prompts- Pick a write-up you actually authored or drove, not one you were merely interviewed for.
- List your root cause and your contributing factors in two separate columns before rehearsing.
- Know how many action items shipped, and prepare the honest story of one that did not.
- This can be the same incident as your broke-production story, re-angled onto the document.
draft and rehearse your own answer in a learn session
go deeper
This probes analytical honesty and follow-through rather than composure in the moment. The interviewer wants to see whether you can hold your own error still on the page, separate mechanism from conditions, and convert analysis into commitments that someone actually owns. A strong answer proves the write-up changed behaviour, and reports the unfinished items as frankly as the finished ones.
The incident was mine. A cooldown value I tuned to stop node flapping ended up applying to a pool class it was never meant to touch, and scale-up on that class stalled long enough that the shared queue backed up through a morning peak. I wrote the write-up the next day and published it in the repository, which is the norm on that project. It opens with the timeline and one plain line saying the change was mine, and then deliberately says nothing more about that, because a document that reads as an apology stops other people volunteering facts. The cause was a label selector dropped during a rebase, so the value matched every pool instead of one. I listed the conditions separately and did not blend them in: nobody who reviewed the patch owned that file, the documented default in our reference page was stale, and we alerted on failed jobs rather than on autoscale delay itself, which is why a downstream project told us before our own dashboards did. We are eleven maintainers across a lot of time zones, so an unowned list would have died. Seven actions, each with a named maintainer and a target release. Five shipped within eighteen days, including the delay alert and a check that fails when a selector matches more pools than the change claims to touch. Two I cut back in the open rather than let them sit. The next selector mistake was caught in review, by that check.
The strength is the analytical split held visibly apart, cause in one sentence and three named conditions beside it, plus honest accounting on what shipped and what was renegotiated publicly. The closing proof is that a later mistake was caught by the artefact. Presenting all seven items as complete would invite a probe this answer would not survive.
My own incident is the one that changed how the project handles all of them. I had shipped a scheduler change that pushed scale-up past four minutes for most of a working day, and while writing it up I noticed I was inventing the format from scratch, exactly as the three authors before me had. So the document had two halves. The first was ordinary: my change, the mechanism, the conditions, the fixes. The second argued that our review process could not have caught this class of bug for anybody, not just for me, because no file in the scaling path had an owner of record and we had no shared definition of what degraded even meant. Out of that I proposed three things to the maintainer group. A template that forces contributing factors to be listed apart from the cause. A published expectation for queue wait, so an incident gets declared by a number rather than by whoever is most annoyed. And a rotating reviewer for write-ups, so the author is never their own only editor. All three were adopted, and I edited the first four reviews myself so the norm was not just a file in the repository. The number I care about is not from my incident. It is that write-ups now close 23 of every 28 actions inside a release cycle, where most used to stay open, and new contributors reach for the template without being told.
This works because the incident is treated as a sample of a class and the output is a durable practice with named components. Doing the first four reviews personally is the detail that separates an adopted norm from a proposal. Without adoption evidence and the follow-through figure, this would read as a suggestion rather than a change.
A contribution counts. Describe the timeline you reconstructed or the section you drafted, and show you understand why the document names systems rather than people while still recording exactly what your change did.
You wrote it for your own incident. Show the analysis: the mechanism, two or three conditions that let it reach production, and at least one action item you personally carried to completion rather than filed.
Show the write-up driving change across a team: actions with named owners and target dates, honest reporting on what slipped, and renegotiating scope in the open instead of letting items rot in a tracker.
Talk about the practice, not one document. Show the template, the review norm or the declaration threshold you introduced after your own incident, who adopted it, and the measured difference in follow-through afterwards.
saying these in an interview costs you the question
- A document that is mostly apology, so the facts never arrive
- Root cause and contributing factors mashed into one undifferentiated list
- Action items with no owner, no date and no follow-up story
- Claiming every item shipped, with no example of one that stalled
- Blameless described as a rule against naming the change author's own change
- No sign anyone outside the author ever read or used the document
- Who read that document, and what did they do with it?Name the actual audience and one concrete consequence: a review that changed, a config someone else fixed, a question a newcomer stopped having to ask. A postmortem nobody used is an essay. If readership was thin, say so and explain what you changed about distribution or format the next time.
- How many of the action items actually shipped?Give the fraction and be honest about the shortfall. Interviewers expect slippage; they are testing whether you tracked it. Say which items you dropped deliberately, and why that was the right trade rather than an oversight. Claiming a perfect completion rate usually invites a probe you cannot survive.
- How do you keep a postmortem blameless when the cause was one person's change?Distinguish naming the change from blaming the author: the document should say exactly what the change did while asking why the system let it through. Note that this is easiest to model when the change is your own, because you can set the register without anyone feeling protected or exposed.
- What did you deliberately leave out of the write-up?Good answers cite speculation you could not evidence, individual performance conversations that belong elsewhere, and rabbit holes that did not contribute. This probe checks editorial judgment; an answer of nothing, we included everything suggests a document long enough that nobody finished it.