What has to be recorded while an incident is still in progress for the postmortem timeline to be usable, and why can't you just reconstruct it afterwards from the chat log?
answer
- observed, done, decided — tag each
- message time is not event time
- the bridge leaves no record
- dashboard links expire, numbers do not
- memory backfills what you now know
basics
~20 sCapture timestamped entries live for observations, actions and decisions, tagged as which of the three they are, plus the four anchor times: detection, declaration, mitigation applied, impact ended. Reconstruction fails because chat records when something was mentioned, not when it happened, and memory rewrites what people knew.
solid answer
~50 sYou need a running log where each entry has a UTC timestamp and is explicitly one of three kinds: something **observed** ("error rate hit 30% at 14:02"), something **done** ("rolled back release 4.2 at 14:31"), or something **decided** ("chose rollback over fix-forward because the change was small and reversible"). On top of that, four anchor timestamps give you the durations everyone will ask about afterwards: detection, declaration, mitigation applied, impact ended — those yield time-to-detect, time-to-mitigate and total impact duration. Reconstruction after the fact fails for concrete reasons, not lazy ones: chat timestamps mark when someone mentioned an action, not when it ran; decisions get made on a voice bridge and leave no trace; hypotheses read as facts in scrollback; dashboard links expire or silently change their default time range; and human memory reliably reconstructs what people knew *then* using what they know *now*, which is precisely the information the postmortem needs unpolluted.
code
json · 8 lines{
"t": "2026-04-11T14:31:07Z",
"kind": "decided",
"who": "ic",
"text": "Roll back release 4.2 rather than fix forward",
"why": "Change was small and reversible; cause still unconfirmed at 29 min",
"evidence": "error_rate 31% since 14:02; deploy 4.2 completed 13:58"
}go deeper
Know that entries need UTC timestamps and that you write down what you did as you do it, not afterwards. Say plainly that the chat log alone is not a timeline.
Explain the observed/done/decided distinction and the anchor timestamps that yield time-to-detect and time-to-mitigate, and give concrete reasons scrollback reconstruction produces wrong durations.
Show you know the timeline is evidence for judging what was reasonable with the information available at the time, and that decisions with their rationale are the entries most often missing and most costly to lose.
Own making capture nearly free: automated deploy and rollback events posting into the incident channel, a recorder who is not on the keyboard, and a mandatory short timeline sweep before responders disperse.
## What the timeline is for The postmortem timeline is not a narrative for readers who missed the incident. It is evidence, and it supports two things that cannot be produced any other way. First, the **durations**: how long until anyone knew, how long until a human was engaged, how long until impact stopped. Those numbers drive whether the follow-up work should target detection, response or the failure itself — three completely different investments. Second, the **decision record**: what each responder believed at the moment they acted. Analysis of contributing factors depends entirely on comparing what was reasonable given the information available at the time against what turned out to be true, and that comparison is impossible once the timeline has been reconstructed by people who now know the answer. ## The three kinds of entry Tag every entry as one of: - **Observed** — a fact from a system: a metric value, an alert firing, a log line, a customer report. Include where it came from so it can be checked later. - **Done** — an action taken against production, with who took it. "Restarted the three workers in eu-west-1" is a timeline entry; "maybe we should restart the workers" is not. - **Decided** — a choice between options, with the reason and the information available at the time. This is the highest-value and most frequently missing category. The distinction matters because scrollback flattens all three into undifferentiated text. Six weeks later nobody can tell whether "looks like the cache is cold" was a confirmed measurement or a guess someone typed while scrolling a dashboard — and in most postmortems that read badly, a guess had quietly become the accepted account of events. ## The four anchors Always pin down: when the failure actually began (often earlier than detection, established later from data), when it was detected, when a human acknowledged and declared, when mitigation was applied, and when impact ended and was confirmed ended. These give the durations that make incidents comparable across a quarter, and they are the ones people argue about afterwards if nobody wrote them down at the time. ## Why chat scrollback is not a timeline Every one of these is a real, repeated failure: - **Message time is not event time.** "Rolled back" typed at 14:36 might describe a rollback started at 14:29 and finished at 14:34. Duration arithmetic built on message timestamps is simply wrong. - **The bridge is silent.** The most consequential decisions in a fast incident are made on a voice call. If nobody types "decision: failing over to the secondary, accepting up to 30 seconds of writes lost", that decision never existed as far as the record is concerned. - **Side channels.** Direct messages, a vendor's support portal, another team's channel and someone's terminal history each hold pieces the incident channel never saw. - **Evidence rots.** Dashboard links resolve to a relative time range and show something different next week; short-retention logs age out; the deployment that caused it gets superseded. Screenshot and paste the numbers into the log while they exist. - **Memory is reconstructive.** People genuinely and honestly remember having suspected the true cause earlier than they did. This is not dishonesty; it is how memory works, and it is exactly the corruption that makes a late-built timeline useless for judging whether the response was reasonable. ## Keeping it cheap The honest tension is that every minute spent recording is a minute not spent fixing, so capture has to be nearly free. What works in practice: someone who is not on the keyboard owns the log (in a structured response this is a dedicated scribe); the incident channel is the single working surface so mentions can be marked into the timeline with a reaction or a bot command; deployment, rollback and configuration systems emit their own timestamped events into the channel automatically so the highest-value entries need no human; and after mitigation, while everyone is still on the call, the group spends ten minutes filling gaps before dispersing. That last habit — a short timeline sweep immediately after impact ends — recovers more than any attempt a week later. ## The trade-off to state out loud You can run an incident with nobody recording and go faster in the first twenty minutes. What you buy with that speed is a postmortem that argues about what happened instead of what to change, and follow-up work aimed at the wrong stage of the response. The cost of live capture is one person's attention; the cost of skipping it lands weeks later on everyone.
- Which single timeline entry is most often missing, and why does its absence hurt the postmortem most?The decision with its rationale. Actions usually leave a trace in a deploy or change system, but why someone chose rollback over fix-forward exists only if a human writes it down. Without it the postmortem cannot tell a reasonable call made under uncertainty from a careless one, which is the difference between learning something and assigning blame.
- How do you establish when the failure actually started, as opposed to when it was detected?From data, after the fact: work backwards through metrics, logs and change records to the first anomalous point, and record it as a separate anchor from detection. The gap between the two is the detection deficit, and it is the number that justifies investment in alerting rather than in the failing component itself.
- Responders are on a voice bridge and moving fast. How do you get their decisions into the record without slowing them down?Give the job to someone not debugging, and have them narrate back into the channel in one line per decision — the incident commander confirming "logging that we are failing over now" takes three seconds and makes the entry authoritative. Automated deploy, rollback and flag-change events posting themselves covers most of the action entries for free.
saying these in an interview costs you the question
- Assuming the chat log is already the timeline
- Recording actions but never the reasoning behind decisions
- Using message timestamps as the time an action actually ran
- Linking a dashboard instead of capturing the numbers
- Building the timeline days later from responders' recollections