Your public API has been returning errors for twelve minutes and the cause is still unknown. How does the public status-page post differ from the internal stakeholder update, and what determines when you post publicly?
answer
- two audiences, two costs of being wrong
- impact confirmed, not cause confirmed
- threshold agreed before the incident
- first post can be refined, never retracted
- concealment costs more than the outage
basics
~20 sPublic posts are triggered by confirmed customer-visible impact, not by a known cause. They state the symptom, scope and next update time in plain language with no internal names or speculation, while the internal update stays candid about hypotheses, ruled-out theories and mitigation options.
solid answer
~50 sThey are two documents for two audiences with different consequences. The internal update can be candid — leading hypothesis, what has been ruled out, which mitigation is being weighed — because the readers are colleagues who understand uncertainty. The public post is a durable record customers will screenshot, so it carries only what is confirmed: the user-visible symptom, its scope, when it started, a workaround if one exists, and when the next post lands. No internal service names, no cause attribution, no blaming an upstream provider. The trigger for posting is **confirmed customer-visible impact**, not diagnosis — you should have a pre-agreed threshold, such as any confirmed impact at your top two severities, so nobody is debating it mid-incident. Twelve minutes of errors with an unknown cause clears that bar: you post "investigating elevated error rates" now and refine later.
go deeper
Know that the public post uses plain customer-facing language and states only confirmed symptoms, while internal updates can discuss suspicions. Say that you would not name internal systems publicly.
Explain that the trigger for publishing is confirmed customer impact rather than a known cause, and describe the fields of a first post: symptom, scope, start time, workaround, next update.
Weigh both costs out loud — a wrong public scope line versus customers learning of the outage elsewhere — and defend the low-commitment first post that can be refined. Show you know a premature resolved-state is worse than a slow update.
Own the policy: a written severity-linked posting threshold and a component-state convention that nobody negotiates during an incident, plus the standing decision about when a public incident report follows and who signs it off.
## Two documents, two audiences, two costs The internal stakeholder update and the public status post look similar and are governed by completely different rules, because the cost of being wrong differs. Internally, a wrong hypothesis costs a correction in the next update among people who expect uncertainty. Publicly, a wrong statement is a durable artifact: it is screenshotted, quoted back to your account team, and occasionally picked up by press. That asymmetry drives every difference between the two. **Internal** can and should be candid: the leading theory, what has been eliminated, which mitigations are on the table and what each risks, the severity and whether it is moving. Responders and leadership need the reasoning, not just the conclusion. **Public** carries only confirmed, user-observable facts: - what the customer sees ("elevated error rates on API requests"), - who is affected — region, product surface, percentage if you can defend it, - when it started, in a stated timezone (UTC is the convention), - a workaround if one genuinely exists, - when the next post will appear. And it deliberately omits: internal service names, unconfirmed causes, the name of any vendor or person, apologies that concede more than you know, and anything that reads as a fix ETA. ## What triggers the first post The trigger is **confirmed customer-visible impact**, not a confirmed cause. This is the crux of the question, because the intuitive instinct — wait until we can explain it — is the wrong one. Customers already know something is broken; they experienced it before your monitoring did in more cases than anyone likes to admit. The first post exists to say "we see it too", which is a different and much cheaper claim than "here is why". The threshold should be written down in advance and tied to severity, for example: any confirmed customer-visible impact at your top two severity levels gets a public post within a fixed number of minutes of declaration. Deciding this in the moment guarantees a slow, argued, politically-shaded answer. Twelve minutes of unexplained API errors clears any sane threshold, so the answer to the scenario is that you should already have posted. ## The cost of each mistake, honestly Posting early has a real cost. You are publishing an admission of an outage before you have confirmed its shape; the scope line may be wrong and need correcting; support volume shifts as customers who had not noticed now check the page; and for some businesses a public post has contractual or commercial visibility. This is why "post early" is a judgment, not a reflex — but the judgment almost always lands on the same side. Not posting costs more. Customers who cannot tell whether the problem is yours or theirs open tickets, which floods support at the worst possible moment; someone posts on social media before you do and now you are responding rather than informing; and the eventual disclosure carries the additional accusation that you concealed it. The trust damage from perceived concealment routinely exceeds the damage from the outage itself. The practical resolution is a low-commitment first post: "We are investigating elevated error rates affecting API requests. We will update by 15:00 UTC." That is honest at twelve minutes with no diagnosis, and it can only be refined, never retracted. ## Cadence and the component model Public cadence is typically slower than internal — a public post every 30 to 60 minutes against a tighter internal loop — but the same rule applies: the post names its own next update time and you honour it. Most status pages model the service as components with a state such as degraded performance versus a full outage; choose the state that matches what a customer experiences, and resist the pull to under-call it. Marking a total failure as "degraded performance" is the most common way a status page loses its credibility, and once customers stop believing the page they go straight to support instead. ## Closing publicly The resolution post says impact has ended and when, and explicitly does not claim the cause is understood if it is not. Whether you publish a public incident report afterwards is a separate, deliberate decision — for a significant outage a summary is worth far more than silence, but it is written after the internal postmortem, not improvised in the resolution post. Never close a public incident before the SLIs actually confirm recovery; a premature "resolved" followed by a reopen is one of the few things that damages credibility more than the outage.
- The outage is caused by a failure at your cloud provider. Should the public post say so?Not as the headline and not as an excuse. Customers bought your service, not your provider's, so the impact and your response are what belong on the page; you may factually reference an upstream provider incident once it is publicly confirmed by them. Attributing blame before confirmation is a retraction waiting to happen and reads as deflection either way.
- Only a handful of enterprise customers are affected. Is a public status post still the right channel?Often not. A narrow, identifiable blast radius is better served by direct notification through account teams, which reaches exactly the affected parties with specific detail. Reserve the public page for impact a customer cannot attribute without it. The decision hinges on whether unaffected customers would be alarmed for no reason.
- How do you keep the status page from itself failing during your outage?Host it outside the infrastructure it reports on — a separate provider, separate DNS, separate authentication path — and make sure the publishing workflow does not depend on your own identity provider. A status page that shares a failure domain with the product goes dark exactly when it is needed, which is a well-known and entirely avoidable trap.
saying these in an interview costs you the question
- Waiting for root cause before posting anything publicly
- Under-calling a full outage as degraded performance
- Naming an upstream cloud provider as the cause on the public page
- Copying internal service names and hypotheses onto the status page
- Marking the incident resolved before SLIs confirm recovery