skip to content

During a long revenue-affecting outage, executives and account managers start messaging responders directly for updates. How would you structure incident communications so leadership stays informed without interrupting the people fixing it?

level: principalimportance: nice to knowfreq 32%

answer

  1. the direct message is a design failure
  2. three surfaces, one source of truth
  3. briefer is not the incident commander
  4. thresholds agreed while nobody is stressed
  5. bring costed options, not a situation

basics

~20 s

Give each audience its own surface with one publishing point: a responder working channel, a broadcast status channel on a fixed cadence, and a named human who briefs leadership directly. Agree notification thresholds by severity in advance so nobody negotiates access during the incident.

solid answer

~50 s

Direct messages to responders are a symptom: leadership has a real need and no channel that serves it, so they take the fastest path, which is also the most expensive one. The structure that fixes it has three layers — the responder channel where work happens, a broadcast status channel carrying updates on a stated cadence, and a named person who briefs executives on a slower loop and absorbs their questions. All three are fed from one publishing point so nobody hears different numbers. Crucially, the thresholds are pre-agreed: which severities notify which leaders, how fast, and through what channel, written down before the incident, because negotiating access mid-outage always resolves in favour of the most senior person asking. The failure mode on the other side is shielding too hard — leadership owns decisions responders cannot make, such as authorising emergency vendor spend, triggering contractual customer notification, or accepting data loss to restore faster. Those need escalating early, with costed options rather than a situation report.

go deeper

for a junior

Know that status requests belong in a stakeholder channel, not the responder channel, and that you can politely point someone there instead of answering mid-incident.

for a middle

Describe the separation between a working channel and a broadcast status channel with a stated cadence, and explain why a single publishing point prevents conflicting impact numbers.

for a senior

Show you would take a named person out of the response to brief leadership and defend that cost, and identify the decisions — emergency spend, customer notification, accepting data loss — that must go up rather than be absorbed.

for a principal

Own the written, severity-linked notification policy agreed in calm conditions, including what is asked of leadership in return, and the habit of escalating business decisions the moment they become plausible with costed options attached.

## Why the direct messages happen An executive messaging a responder at hour two of a revenue-affecting outage is behaving rationally. They have accountability, an unanswered question, and no channel that reliably answers it. The instinct to treat this as a discipline problem — "leadership should know better" — is wrong and does not work. It is a design problem: build the channel that serves the need better than a direct message does, and the direct messages mostly stop. The cost being avoided is specific. Interrupting a responder mid-diagnosis does not cost the thirty seconds of the reply; it costs the reconstruction of everything they were holding in their head, and a senior person's question cannot be deferred the way a peer's can. Multiply by several leaders across a multi-hour incident and it is a substantial share of your response capacity. ## Three surfaces, one publishing point **Responder working channel.** Where diagnosis and coordination happen. Norm: if you are not actively responding, you read but do not ask. **Stakeholder status channel.** Broadcast-only, updates on a stated cadence, written in impact terms. This absorbs most of the demand, because most people who ask for status simply want status. **Executive briefing.** A slower loop — perhaps hourly or on material change — carrying what leadership specifically needs: customer and revenue impact, whether it is contained, what decisions are pending on them, and the next decision point. Delivered by a named human who can also take questions, because a broadcast feed alone does not satisfy someone who is accountable and wants to ask one thing. All three derive from one source of truth. The moment two people are independently drafting updates you get two versions of the impact number circulating, and an executive quoting the stale one to a customer. The person doing the briefing must not be the incident commander. Command capacity is exactly what a long incident runs out of. In a structured response this is the communications role; in a small team it is a manager or a second engineer deliberately taken out of the debugging. The trade is explicit and worth stating in an interview: you spend one person's time to protect the rest of the response, and at high severity that is unambiguously worth it. ## Pre-agreed thresholds The part that fails in practice is not the channel design but the escalation policy behind it. Write down, before any incident: which severities notify which leaders, within how long, through which mechanism, and who is authorised to make that call. If this is not settled in advance it gets negotiated during the fire, and the negotiation is always won by the most senior person asking — which is how you end up with three vice presidents in the responder channel. The same document should say what is expected of the leaders it notifies: read the status channel, route questions to the named briefer, do not direct responders. Leadership generally honours that when it is a written norm agreed in calm conditions and asked of everyone, and generally does not when it is asserted for the first time by a stressed engineer at 03:00. ## What leadership is genuinely for Shielding is a means, not a goal, and over-shielding is its own failure. Some decisions are outside a responder's authority and belong to leadership, and they take real time to make: - authorising emergency spend — premium vendor support, an unplanned capacity purchase; - triggering contractual or regulatory customer notification, which may run on its own clock; - accepting a business trade-off, such as restoring service faster while losing some recent data, or disabling a revenue-generating feature to stabilise everything else; - pausing an adjacent launch or campaign that would add load; - deciding what is said publicly when the outage becomes a press question. Escalate these **when they become plausible, not when they become necessary**, because the decision-maker may need to be found, briefed and given options. And bring them as options, not as a situation: two paths with their costs and a recommendation gets an answer in minutes, while an undifferentiated status report gets a meeting you have to attend. ## Long incidents and handover Over many hours the comms role needs relief on the same rotation as everyone else, and the handover must transfer the audience list, what has already been promised, the exact cadence committed to, and the outstanding questions. A missed executive update after a shift change reads as the response falling apart, regardless of what is actually happening on the technical side. ## The judgment being tested There is no universally right amount of leadership involvement, which is why this sits at a principal level. Too little and you lose access to the decisions and the budget only they control, and you surprise people whose accountability makes surprise expensive. Too much and you convert responders into reporters. The defensible position is a written, severity-linked policy that gives leadership a genuinely better path than the direct message, plus a standing habit of escalating business decisions early — and then holding the line on the working channel because the norm was agreed in advance rather than invented under stress.

  • An executive insists on joining the responder channel during a major outage. Do you refuse?
    No — refusing access to an accountable leader is a fight you lose and should not pick. Set the terms instead: they are welcome to observe, questions go to the named briefer, and no instructions to responders. That is enforceable when it is a pre-agreed norm applied to everyone, and it converts a disruptive presence into a well-informed one.
  • What makes an executive briefing different in content from the general stakeholder update?
    It is decision-oriented rather than descriptive. Customer and revenue impact, whether the impact is contained or still growing, what is pending on them right now, and the next point at which a call is needed. The general update answers "what is happening"; the executive briefing answers "what do you need from me and when".
  • How do you keep leadership from over-reacting to an early impact estimate that later proves wrong?
    State confidence explicitly with every number and say what would change it: "roughly 30% of requests, measured at the edge, unconfirmed for the mobile clients — we will know by 15:00". Numbers offered without a confidence band get treated as facts, quoted to customers, and are then very expensive to revise downward.

saying these in an interview costs you the question

  • Treating executive interruptions as a discipline problem, not a design gap
  • Letting the incident commander also brief leadership
  • Shielding leadership from decisions only they can authorise
  • Negotiating who gets notified in the middle of the outage
  • Two people drafting updates and circulating different impact numbers

context