skip to content

You are handling communications for an ongoing SEV1 outage where customer logins are failing. What goes into each stakeholder update, how often do you send one, and what should you never promise in it?

level: middleimportance: must knowfreq 70%

answer

  1. cadence agreed before the fire
  2. impact stated in user terms
  3. commit to the next update time
  4. an ETA is a promise you cannot keep
  5. silence reads as nobody working

basics

~20 s

Incident updates go out on a fixed cadence tied to severity — commonly every 30 minutes at the top severity — and each one states user-visible impact, what the response is doing now, and the time of the next update. Never promise a restoration ETA.

solid answer

~50 s

Every update carries three things: the impact in user terms ("roughly 30% of login attempts have been failing since 14:02 UTC"), the current state of the response including what has been ruled out, and an explicit **next update time**. The cadence is decided before the incident and tied to severity — say 30 minutes for a SEV1, hourly for a SEV2 — plus an immediate out-of-band update on any material change: severity moving, a mitigation going in, impact ending. You send the update even when nothing has changed, because silence reads as "nobody is working on it". The one thing you never give is a restoration ETA. You almost never know it; support will relay it, customers will plan around it, and when it slips you have manufactured a second trust problem on top of the outage. Commit to the next update time instead — that is the one thing you fully control.

go deeper

for a junior

Know that updates go out on a fixed schedule, not when someone asks, and that each one names the user-visible impact and the time of the next update. Say plainly that you would not give a fix ETA.

for a middle

Explain the three fields of an update and why the next-update time replaces an ETA, and describe tying cadence to severity with immediate out-of-band updates on material changes.

for a senior

Show judgment about what leaves the responder channel: separating confirmed facts from hypotheses, refusing to name an unconfirmed cause, and splitting "impact ended" from "cause understood" in the closing update.

for a principal

Own the trade-off that comms cost responder attention. Be ready to argue for cadences and templates fixed in policy per severity, and for a writer who is not on the keyboard, rather than leaving it to whoever is closest to the problem.

## What an incident update is actually for A stakeholder update is not a debugging log and not a status summary for your own team. It exists so that three groups can make decisions without interrupting the responders: support and account teams who are answering customers right now, leadership who may need to make a business call, and other engineering teams deciding whether to hold their own releases. Everything in the update should serve one of those decisions. Anything that serves none of them — stack traces, service names nobody outside the team knows, running commentary on a hypothesis — belongs in the responders' working channel. ## The three fields every update carries **Impact, in user terms, with a start time.** Not "the auth service is degraded" but "users cannot sign in; roughly 30% of login attempts have failed since 14:02 UTC; already-signed-in sessions are unaffected". Scope and magnitude matter as much as the symptom, because support answers different questions for "all customers" than for "customers in one region". If you have a workaround, it goes here — it is the single most valuable line in the update for anyone on a phone with a customer. **State of the response.** What is confirmed, what has been ruled out, what is being attempted right now. Label hypotheses explicitly as hypotheses. If a mitigation is in flight, say so and say what you will look at to know whether it worked. This is where a candidate reveals whether they understand the difference between reporting facts and leaking speculation: "we suspect a bad deploy" travels outward, gets repeated as fact, and has to be publicly retracted when it turns out to be a dependency. **The next update time.** This is the field people skip and the one that does the most work. "Next update at 15:00 UTC" converts an anxious audience into a patient one and stops the stream of individual "any news?" messages aimed at the people trying to fix things. Missing a time you committed to is worse than committing to a wider interval, so pick a cadence you can actually sustain. ``` SEV1 — Login failures Impact: ~30% of sign-in attempts failing since 14:02 UTC. Existing sessions unaffected. All regions. Status: Bad configuration rolled out at 13:58 identified as the likely cause (unconfirmed). Rollback started at 14:31. Next update: 15:00 UTC, or sooner if impact changes. ``` ## Cadence Tie the interval to severity and agree it in advance so nobody is negotiating it mid-incident: something like 30 minutes at the top severity, hourly a tier down, and always an immediate update on a material change — severity upgraded or downgraded, mitigation applied, impact ended. The heartbeat rule matters: send the update on schedule even when there is nothing new. "Still investigating, impact unchanged, next update 15:30 UTC" is genuine information. It says the response is alive and staffed. A gap says the opposite, and people fill silence with their own worst guess. ## What you never promise A restoration ETA. During an incident you rarely know the cause, so you cannot know the fix time; an ETA is a guess wearing the costume of a commitment. It propagates fast — support tells customers, customers tell their own users, someone builds a plan around it — and when it slips you now have two problems. The professional substitute is a next-update time plus, if you genuinely have one, a bounded statement about a specific action: "the rollback completes in about ten minutes; we will confirm recovery in the 15:00 update." Similarly, do not name a cause you have not confirmed, and never name a person or a vendor as the reason. ## Closing the incident out The final update must separate two things that stakeholders routinely conflate: **impact has ended** and **we understand why it happened**. Declaring recovery is a statement about the SLIs you are watching now; the explanation comes later. Say plainly that a postmortem will follow, and do not improvise a prevention plan in the resolution message. ## The cost side Updates are not free. Each one consumes a responder's attention at exactly the moment attention is scarcest, which is precisely why the cadence, the audiences and the template are settled before the incident rather than during it, and why the person writing them should not be the person deep in the debugger. That trade — a fixed, slightly generous cadence written by someone not on the keyboard, versus ad-hoc updates from whoever is closest to the problem — is the real decision behind the question.

  • It is the 30-minute mark and the investigation has produced nothing new. Do you skip the update?
    No — send the heartbeat. "No change in impact, still investigating, next update at 15:30 UTC" tells stakeholders the response is staffed and progressing. A missed update is read as abandonment, and it reliably produces exactly the direct messages to responders that the cadence exists to prevent.
  • A mitigation has worked and errors are gone, but you have no idea what caused it. What does the update say?
    Declare mitigation, not resolution: state that impact ended at a specific time, name the SLI you are watching to confirm it holds, and say the cause is not yet understood and a postmortem will follow. Conflating "errors stopped" with "we fixed it" is what produces a public retraction when the problem recurs an hour later.
  • A vice president asks in the channel for a fix ETA and will not accept "next update at 15:00". How do you answer?
    Give what you actually know instead: the specific action in flight, how long that action takes, and when you will know whether it worked. "Rollback finishes in ten minutes; we confirm recovery at 15:00" is bounded and honest. Escalate the pressure itself to the incident commander rather than inventing a number to end the conversation.

saying these in an interview costs you the question

  • Waiting until root cause is known before communicating anything
  • Promising a restoration ETA under pressure from stakeholders
  • Describing impact by internal service name instead of user symptom
  • Skipping the scheduled update because there is nothing new
  • Sending stakeholders raw debugging chatter from the responder channel

context