Rolling back, restarting, or replacing instances during an incident usually destroys the state you would need to explain the failure later. What do you capture before you pull the mitigation lever, and how do you keep that capture from delaying the mitigation itself?
answer
- what dies with the process
- central telemetry survives a rollback
- keep one sick instance out of rotation
- seconds, not minutes, of capture
- the timeline is evidence too
basics
~20 sCapture only what dies with the process — thread stacks, heap state, local files, in-memory queues — and prefer quarantining one failing instance out of rotation over restarting them all. Anything already shipped to central logging, metrics or tracing survives the mitigation, so do not wait for it.
solid answer
~50 sI split the evidence into what survives the mitigation and what does not. Logs, metrics and traces already shipped centrally survive a rollback or restart, so I never delay for them. What dies with the process is in-process and node-local state: thread dumps, heap contents, in-flight queues, files on the instance's disk. The cheapest trick is a **specimen** — pull one failing instance out of the load-balancer rotation and leave it running instead of restarting the whole fleet. It serves no users, so it costs nothing, and it preserves a live, poke-able copy of the failure. The rule that keeps this honest is a timebox: if a capture takes seconds and needs no thought, take it; if it needs more than about a minute or any debate, skip it and mitigate. User impact always outweighs postmortem quality. I also screenshot or note the current dashboard state and the exact times of every action, because that timeline is the part people reconstruct badly afterwards.
go deeper
Know that restarting or rolling back throws away the state that explains the failure, and that a quick thread dump before you do it costs seconds. Do not let evidence collection stop you from mitigating.
Be able to sort evidence into what survives the mitigation and what does not, and name the ephemeral list: thread stacks, heap, in-memory queues, node-local files. Explain the quarantine-one-instance trick and what it costs.
Show the timebox judgment out loud — seconds and no thinking, or skip it — and say when the specimen must be killed instead of kept because it is causing the impact. Mention snapshotting graphs and timestamping every action.
Frame this as an infrastructure decision rather than an incident decision: continuous profiling, automatic dumps, retained previous-generation instances and durable telemetry mean responders never have to trade user impact against explainability in the first place.
## Why this is a real tension, not a platitude Mitigation-first response has a genuine cost, and this is it. The fastest mitigations — restart, roll back, replace the instances, fail away from the region — all work by *destroying the state that is misbehaving*. That is precisely why they work, and it is also why the postmortem sometimes opens with "we never found out what happened." A candidate who has actually carried a pager knows this tradeoff exists and has a cheap answer to it. A candidate who has only read about incidents either forgets evidence entirely or, worse, spends fifteen minutes collecting it while users fail. ## Sort the evidence by whether it survives The useful mental split is durable versus ephemeral. **Durable — do not wait for it.** Anything that has already left the process is safe: log lines shipped to a central store, metrics scraped or pushed to a time-series backend, traces exported to a tracing backend, load-balancer and proxy access logs, audit records of deploys and configuration changes. All of it will still be there after the rollback. Collecting it during the incident buys nothing and costs minutes. **Ephemeral — dies with the process or the node.** This is the list worth spending seconds on: - **Thread stacks.** What are the threads actually blocked on? This is often the whole answer for a hang or a deadlock, and it is unrecoverable once the process exits. - **Heap or memory state.** For a suspected leak or an out-of-memory condition, a heap snapshot is the only artifact that shows what was actually retained. It is also the most expensive capture on this list, and may itself stall the process — take it from an instance that is already out of rotation. - **In-memory and local queues.** Depth, oldest entry, and whether anything is being dropped. - **Node-local files.** Crash artifacts, core dumps, or logs that were never shipped because the shipping agent was itself failing. - **Live connection and socket state.** Whether the process is holding connections open to a dependency that has stopped answering. ## The specimen: the highest-value habit here The technique that resolves the tension almost entirely is **quarantine one, mitigate the rest**. Remove a single affected instance from the load-balancer or service-mesh rotation, or scale the fleet by one and leave the sick one running, then roll back or restart everything else. Users are recovered on the healthy fleet within the normal mitigation time, and you keep a live, still-broken specimen you can attach a debugger to, dump repeatedly, and reason about at leisure. This costs one instance's worth of capacity and roughly the same number of seconds as the mitigation itself. Note the constraint that makes it fail: if the failure is *caused* by the specimen — it holds a distributed lock, it is poisoning a shared cache, it is hammering a dependency — quarantining it does not stop the impact and you must kill it. Judgment, not ritual. ## The timebox is the actual rule Evidence collection expands to fill the incident if you let it. The bound that experienced responders apply is roughly: **if it takes seconds and requires no thinking, do it; otherwise mitigate now and accept the loss.** Concretely, a thread dump on one instance is fine. Deciding which of four possible diagnostic tools to attach, and to which instance, is not — that is diagnosis wearing an evidence-collection costume. The way to make more evidence survivable is to move that work out of the incident entirely: continuous profiling, automatic heap dumps on out-of-memory conditions, log shipping that keeps up under load, and retaining the previous generation of instances briefly after replacement. Every one of those turns an in-incident decision into a decision you already made, calmly, months ago. ## The evidence nobody thinks of The most commonly lost artifact is not a dump. It is **the timeline**: what was observed at what time, what action each responder took, at what time, and what happened next. Human memory reconstructs this badly and confidently. Two habits fix it: announce every action in the incident channel *before* taking it with a timestamp, and snapshot the graphs you are looking at now — dashboards are frequently rebuilt, and retention windows and downsampling can quietly coarsen the exact minutes that mattered by the time someone opens the postmortem a week later.
- When would you refuse to keep a failing instance around as a specimen?When the instance is itself part of the impact. If it holds a lock other replicas are waiting on, is writing corrupt data, is poisoning a shared cache, or is retrying hard enough to keep a dependency down, then leaving it alive prolongs the incident. In that case kill it and settle for whatever dumps you took first.
- How do you make more evidence survive future incidents without slowing anyone down?Move the collection out of the incident. Continuous profiling, automatic heap dumps on out-of-memory, log and trace shipping that keeps up under load, and briefly retaining the previous generation of instances after replacement all mean the artifact already exists when you mitigate. The best in-incident evidence decision is one that was made months earlier.
- Everything is centrally logged. Is there still anything worth capturing before a rollback?Yes — thread stacks and memory state are never in your logs, and they are what explain hangs, deadlocks and leaks. Node-local artifacts matter too when the shipping agent itself was failing, which is common under exactly the load conditions that cause the incident. Central telemetry tells you what happened; the in-process state often tells you why.
saying these in an interview costs you the question
- Spends fifteen minutes collecting dumps while users fail
- Restarts the whole fleet with no specimen kept
- Waits to export logs that are already shipped centrally
- Assumes central dashboards will still show that minute later
- Treats evidence capture as optional every single time