How do you reconstruct the timeline of an incident for its postmortem, and which timestamps in that timeline matter most?
answer
- evidence first, memory last
- normalize everything to UTC
- impact started before you noticed
- the gaps, not the total duration
- record what they believed, not just what they did
basics
~20 sRebuild the timeline from machine evidence — deploy records, alert history, chat transcripts, dashboards — normalized to UTC. The load-bearing timestamps are impact start, detection, human engagement, first mitigation and impact end, because their gaps measure detection and mitigation separately.
solid answer
~50 sI build it from evidence, not memory. Deploy and change records, alert firing times, the incident chat transcript, dashboard screenshots and command history all get merged into one list normalized to UTC, because a timeline stitched from three people's recollections in three time zones is where postmortems quietly go wrong. Five timestamps carry the analysis: when impact actually started (usually earlier than anyone noticed), when monitoring detected it, when a human actually engaged, when the first mitigation was applied, and when impact ended. The gaps between them are the real findings — impact-to-detection is a monitoring problem, detection-to-engagement is a paging and escalation problem, engagement-to-mitigation is a runbook and access problem, and they have completely different fixes. I also record what responders *believed* at each point, not just what they did, because a 25-minute gap explained by "we were looking at the database because the dashboard implicated it" is a finding about the dashboard.
go deeper
Be able to list the evidence sources — deploy records, alerts, chat, dashboards — and say that the timeline is assembled from those in UTC rather than from what people remember.
Explain the five key timestamps and what each interval measures, and show that detect-time and mitigate-time are separate problems with separate fixes.
Demonstrate that you capture responder belief and available information at each step, and use the interval breakdown to argue for a specific investment rather than a generic "improve monitoring".
Own the instrumentation of the process itself: scribing during the incident, retention long enough to reconstruct, and consistent interval definitions so durations can be compared across incidents and quarters.
## Why the timeline is the load-bearing section Every other part of a postmortem is an interpretation. The timeline is supposed to be the record. If it is wrong, the analysis built on it is wrong in ways nobody can see later, and the postmortem becomes a story rather than evidence. It is also the only section that produces numbers, and the numbers are what make the findings arguable rather than rhetorical. ## Sources, in rough order of trustworthiness 1. **Change and deploy records.** Exact, machine-generated, and usually the closest thing to a trigger timestamp. 2. **Alerting history.** When each alert fired, re-fired and resolved. This gives detection time and also shows which alerts fired and were ignored. 3. **The incident chat transcript.** The single highest-value human source, because it is contemporaneous — it captures what people believed *at the time*, before the answer was known. 4. **Metrics and logs.** Re-query them rather than trusting a screenshot; you often find impact began before the alert threshold was crossed. 5. **Ticketing and paging system records.** Page sent, page acknowledged, escalated to secondary. 6. **Human recollection.** Last, and always reconciled against the above. Memory of an incident is reliably compressed and reordered. Normalize everything to UTC as you go. Mixed local time zones and unsynchronized laptop clocks are a routine source of timelines that imply a mitigation happened before the change that caused it. ## The five timestamps that carry the analysis - **Impact start (t0).** When users first experienced degradation — not when you noticed. Found by walking the SLI back until it leaves its normal band. - **Detection (t1).** When monitoring produced a signal, whether or not a human saw it. - **Engagement (t2).** When a human acknowledged and began working. - **First mitigation applied (t3).** The action that actually reduced impact, which is often not the first action attempted. - **Impact end (t4).** When the SLI returned to normal. Recovery of the underlying defect can be much later and is a separate mark. What matters is the intervals, because each one indicts a different system: ``` t1 - t0 time to detect -> monitoring coverage, thresholds, SLI choice t2 - t1 time to engage -> paging, routing, escalation, alert fatigue t3 - t2 time to mitigate -> runbooks, access, tooling, diagnosis difficulty t4 - t3 time to recover -> propagation, cache TTLs, restart and drain time ``` A team that reports only "the incident lasted 74 minutes" cannot tell whether to invest in monitoring or in rollback tooling. A team that reports 22 / 3 / 41 / 8 knows immediately that diagnosis was the expensive part. ## Recording belief, not just action For each significant entry, note what the responders thought was happening. "14:22 — restarted the cache tier, believing the latency was cache eviction" is a far more useful line than "14:22 — restarted the cache tier." Long unexplained gaps are usually not idle time; they are time spent on a hypothesis that turned out to be wrong, and *why it was plausible* is the finding. Frequently the answer is that a dashboard, a misleading alert name, or a runbook pointed there. ## Guarding against hindsight Writing the timeline after you know the answer makes every wrong turn look obviously wrong. Two habits help. First, mark clearly what information was available at each timestamp rather than what was true — the signal that would have identified the cause may not have existed on any dashboard yet. Second, write the timeline before writing the causes section, and resist editing entries to fit the conclusion you later reach. ## Practical mechanics Start the timeline during the incident, not after. A scribe pasting timestamped notes into the incident channel costs almost nothing and saves hours of reconstruction. Anything not captured live must be re-derived from logs that may have already rolled off short retention — which is itself a finding worth writing down. ## In an interview Name the sources, name the intervals, and give one example where separating detect-time from mitigate-time changed the investment decision. That is the answer of someone who has actually written one.
- How do you find the true impact-start time when no alert fired until much later?Walk the SLI backwards from the alert until the metric leaves its normal band, and corroborate with a second signal — customer-side errors, support tickets, or an upstream client's latency. If the earliest bad datapoint sits at the edge of retention or scrape granularity, say so explicitly and treat the coarse resolution as its own finding rather than quietly rounding.
- A responder's memory contradicts the chat transcript. Which wins?The transcript, for the record of what happened and when. The memory is still valuable but as a different kind of evidence: it tells you what the experience felt like and what was confusing. I record the transcript timestamp in the timeline and, where the discrepancy is meaningful, note the recollection separately — a responder who remembers being paged much later than the log shows is often reporting a real notification-delivery problem.
- Why keep recovery time separate from mitigation time?Because they buy different things. Mitigation is when user impact stops, which is what the SLO and the customer care about. Recovery is when the system is fully back in its intended state — failover reverted, backfill complete, defect actually fixed. Merging them hides fast mitigation behind slow cleanup and makes rollback tooling look less effective than it is.
saying these in an interview costs you the question
- Starts the timeline at the alert rather than at first user impact
- Builds it from recollection after the fact
- Reports only total duration with no detect and mitigate split
- Mixes local time zones without normalizing
- Edits entries so the wrong turns look avoidable in hindsight