It is 02:10 and you have one ambiguous signal: elevated 5xx errors on one of six API instances, no customer reports yet. Make the case for declaring an incident now on suspicion versus waiting for confirmation, and say how you would make that call repeatable for your team.
answer
- the two errors are priced differently
- undo costs one message
- ramp-up is serial if you start late
- still asking after five minutes?
- zero false declarations means bar too high
basics
~20 sDeclare now. The costs are asymmetric: an over-declaration is undone with one message and a few interrupted minutes, while a late declaration adds unmitigated impact plus the response ramp-up you could have run in parallel. Declare low and adjust.
solid answer
~50 sThe two errors are not symmetrically priced. Declaring wrongly costs perhaps three people times twenty minutes, and it is reversible with a single message. Declaring late costs impact accruing every minute *plus* the ramp-up you now have to do serially — finding the deploy history, pulling in the service owner, opening the channel, getting the runbook — work that would have overlapped the investigation if it had started earlier. So I declare at a low severity, say SEV3, on the suspicion, and adjust once I know more; the grade is cheap to move. To make it repeatable I give the team an explicit tripwire rather than a judgment call: if you are still asking yourself whether this is an incident five minutes in, it is one. I also make closing a declared incident as low-friction as declaring it, and I watch the false-declaration rate — if it is near zero, the bar is too high and real incidents are being handled quietly by one person.
go deeper
Know that you are allowed to declare an incident yourself without asking permission, and that declaring something that turns out to be nothing is a normal, acceptable outcome. Say what you would write in the first message: what you saw, when, and how big it might be.
Explain the cost asymmetry concretely — a reversible interruption versus impact that accrues per minute — and that you would declare low and adjust the grade as facts arrive rather than guess the final severity up front.
Bring the second cost: ramp-up is serial when you declare late and parallel when you declare early. Give the tripwire you actually use, and show you have stood an incident down and treated that as a correct outcome rather than an embarrassment.
Own the incentives and the measurement. Argue why false-declaration rate near zero is a defect signal, how declaration-to-consequence wiring drives hesitation, and where you would deliberately keep a high bar because a declaration triggers external commitments.
## The decision, priced honestly This is an expected-cost question, not a philosophical one. There are two ways to be wrong and they do not cost the same. **Cost of declaring and being wrong.** A handful of people are interrupted. Say three responders for twenty minutes: roughly one engineer-hour, some of it at an unpleasant time. It is fully reversible — one message saying "stood down, monitoring artefact, no impact" — and the record shows a closed non-incident, which is a perfectly respectable entry. **Cost of waiting and being wrong.** Two costs, and the second is the one people forget. The first is direct: impact continues to accrue per minute. If those 5xx errors are one instance in six, roughly 17% of requests may be failing, and every ten minutes of hesitation is ten minutes of that. The second is structural: **response ramp-up is serial if you start it late and parallel if you start it early.** Opening a channel, pulling in the service owner, retrieving the deploy history, finding the runbook, getting a second pair of eyes — that is ten to fifteen minutes of setup. Declared at 02:10, it runs alongside your investigation. Declared at 02:40 because you wanted certainty, it runs *after* it, and you have added it to the outage rather than hidden it inside. That asymmetry is the whole answer: the recoverable error is cheap and bounded, the unrecoverable one grows with time. When the two errors are priced that differently, you take the cheap one. ## What "declaring on suspicion" actually means It does not mean declaring a SEV1 for every twitch. It means separating the *declaration* from the *grade*. Declaring is the act of making the response public and owned; the grade is a dial you can move in either direction as facts arrive. So the correct move at 02:10 is a low-severity declaration with an honest impact statement — "since 02:05, elevated 5xx on api-3; roughly 15% of requests may be affected; cause unknown; investigating" — not a confident guess at the final severity. This also does the thing that pays off most in the following days: the timeline starts recording now, while what you know is what you actually knew. Reconstructing the first thirty minutes afterwards from memory and chat scrollback is the single most common reason a review has a hole in exactly the place the interesting decisions were made. ## Making it repeatable rather than heroic A culture of "use good judgment" produces exactly the variance you are trying to eliminate, because at 2 a.m. the judgment being used is tired and alone. Replace it with mechanisms: - **A tripwire.** "If you are still asking whether this is an incident five minutes after the signal, it is one." The question itself is the trigger; no judgment required. - **Impact-uncertain defaults.** Publish the rule that an unexplained signal with plausible user impact is declared at a specified low tier and revisited within a fixed interval. Removes the choice from the individual. - **Symmetric cost.** Closing must be as easy as opening — one message, no form, no apology, no review obligation for a stood-down non-incident. If closing costs more than declaring, people stop declaring. - **Explicit permission, held by the person with the signal.** The on-call engineer can declare without asking anyone. Any requirement to get approval first inserts a paging round-trip into the decision and guarantees delay. - **Praise the false positive out loud.** The single most effective lever is what happens socially to the person who declared and was wrong. If that is treated as correct behaviour under uncertainty, people keep doing it. ## The measurement that tells you the bar is wrong Track the share of declarations that are stood down with no impact. **A team with near-zero false declarations does not have excellent instincts; it has a bar set too high**, and the tell is that real incidents are being resolved quietly by one person and never recorded. A healthy rate is comfortably non-zero. Track the other direction too: the gap between the first signal and the declaration. If it is routinely tens of minutes on incidents that turned out real, hesitation is a systematic cost you can name and attack. ## The honest counter-argument There is a real limit. Declaring is not free where declaration is wired to heavyweight consequences — a mass page, an executive notification, a public status update, a mandatory review. If that wiring exists, people will rationally hesitate, and the fix is to make the *low* tiers genuinely cheap rather than to ask people to be braver. Similarly, if a specific alert is known-flaky, declaring every time it fires trains the team to ignore declarations; fix the alert instead of absorbing it as a permanent tax on attention.
- Doesn't declaring on suspicion just train people to ignore declarations?Only if the low tiers are wired to heavyweight consequences. If declaring a SEV3 mass-pages a team or triggers a public update, people rationally hesitate and everyone else tunes out. Keep low tiers genuinely cheap — a channel, an owner, a timeline — and reserve broad notification for grades that earn it. And fix known-flaky alerts rather than absorbing them.
- What severity would you actually declare at, given you don't know the impact yet?A low one, with an honest uncertainty statement, and a fixed time to revisit. Separate the declaration from the grade: declaring makes the response public and owned, while the grade is a dial that moves in either direction. Guessing high to be safe distorts the record and burns the credibility of the top tier.
- How would you know your team's declaration bar is set correctly?Two numbers. The share of declarations stood down with no impact — near zero means the bar is too high and real incidents are being handled quietly by one person. And the gap between first signal and declaration on incidents that turned out real; if that is routinely tens of minutes, hesitation is a measurable cost you can attack directly.
A smoke alarm at 2 a.m. is worth getting out of bed for even though it is usually burnt toast; the cost of being wrong is a few cold minutes, and the cost of being right and slow is the house.
saying these in an interview costs you the question
- Wait until you are sure before declaring anything
- Declaring is embarrassing if it turns out to be nothing
- Ask a manager for permission before declaring
- One instance failing is never worth declaring
- No customer complaints means no incident