During an incident, the next fact about a managed tier has to come from a support queue — how does that change how you run it?
answer
- two clocks, you own one
- queue time starts at case open
- first submission should need no reply
- response commitment is not resolution
- mitigation must not need their answer
basics
~20 sYou are now running two clocks and you control only one. Open the case early and in parallel, make it actionable on first read, and drive a mitigation that does not depend on the provider's answer — because the one thing a queue never returns on your schedule is a fix.
solid answer
~50 sOn a machine you operate, the next fact comes from a stack trace and arrives when you look. On a managed tier it comes from someone else's queue and arrives when they get to it, so queue time becomes the long pole and it starts when you open the case, not when you finish investigating. Three habits follow. Open early, in parallel with your own elimination, because the case costs you nothing to have running. Make the first submission answerable — timestamps with a timezone, the resource identifiers, the window of exported metrics, what you already ruled out, and the business impact — or the first reply is a request for information and you have burned a round trip. And pursue a mitigation that works without their answer: shed load, degrade the feature, fail over if you can, route away. Your support plan's response commitment is a promise that someone replies, not a promise that anything is fixed.
go deeper
Recall that when the fault is inside a managed service, you cannot read the failure yourself: you raise a case with the provider and their reply time is not something you control.
Explain why queue time starts when the case is opened, and what a first submission must contain — timestamps, identifiers, visible signals, ruled-out causes, impact — to avoid losing a round trip to a request for information.
Demonstrate running both clocks: case open early and well-formed, a mitigation that works with no reply at all, and stakeholder communication that separates what you can commit to from what you cannot.
Treat the response commitment as a decision made at purchase time, not at incident time, and accept publicly that part of your time-to-resolution now sits in another organisation's queue for every workload placed on a managed tier.
## Two clocks, and you own one An incident on infrastructure you operate has one clock: your investigation. An incident on a **managed tier** has two, because at the moment your own visible evidence is exhausted, the next fact lives with the provider. That second clock has properties yours does not: it starts when the case is opened rather than when you start looking, its pace is not something you can influence much, and it has no relationship to your customer impact except through whatever severity you claimed and whatever plan you bought. Everything in this section follows from one asymmetry: **you can make the second clock start earlier and run better-informed, and you cannot make it run faster.** ## Open early, in parallel The common mistake is treating the case as an admission of defeat — investigate thoroughly, exhaust every avenue, then open it. That sequencing adds your entire investigation time to the front of the queue time. Open it as soon as the store is plausibly implicated, keep investigating, and update or close the case if your own work resolves it. An unnecessary case costs you almost nothing; a case opened an hour late costs you an hour of the only clock you cannot speed up. ## Make the first reply an answer, not a question Every round trip through a queue is expensive, so the first submission has to be answerable without a follow-up. Include: - **Exact timestamps with the timezone**, and the window, not "this morning". - **The identifiers** of the affected resources, unambiguously. - **The signals you can see** and what they show — exported metrics, the engine's own statistics, your client-side latency and pool-wait telemetry. - **What you have ruled out**, explicitly: this is what moves the case past the standard first-response script. - **The impact**, in business terms, which is what the severity claim has to be defensible against. - **The exact question you want answered**, phrased as something only they can see: host-side counters, a fleet-level event, whether a maintenance action touched this instance. A case that says "the database is slow, please help" gets the generic first reply it deserves. A case that says "between these timestamps, client-side latency rose while exported instance metrics stayed flat and session counts were well under the ceiling; we have ruled out queueing and statement-level contention; what do host-side counters show" gets routed differently. ## What the support plan actually commits to A paid support plan typically states a **response-time commitment**: someone will reply within a stated period for a given severity. Read that precisely. | What it is | What it is not | |---|---| | A first-response time for a claimed severity | A time to resolution, or any promise of a fix | | A path to engineers who can see what you cannot | Access for you to that same evidence | | Something you bought before the incident | Something you can upgrade usefully mid-incident | It is also a separate instrument from the service level agreement on the service itself. An availability commitment is remedied with a service credit against the bill when the service missed its published target; a support response commitment is about answering. Conflating the two — expecting a credit because a case was slow, or expecting fast support because availability was breached — is a reliable way to be disappointed by both. ## Meanwhile: mitigate on your own clock The discipline that separates a good incident from a bad one is that the mitigation never depends on the provider's reply. Options that stay yours: 1. **Shed or shape load** — queue the writes, rate-limit the noisiest caller, pause the batch job competing with the ledger. 2. **Degrade deliberately** — turn off the expensive feature, serve a cached or stale view, accept a reduced function that keeps the money moving. 3. **Move the traffic** — route reads to a standby replica that already exists, fail over if the control action is available, drain to a second path if you built one. 4. **Stop the bleeding upstream** — back-pressure at the edge so the pile-up does not turn a slow store into a cascading failure. Note that some of these are themselves control actions that go through the provider's management path, which is precisely why the mitigation plan should include at least one that does not. ## Communicating, and the gap in the timeline afterwards To stakeholders, be exact about which clock you are quoting. You can commit to your mitigation and its time; you cannot commit to their fix, and inventing an estimate on their behalf is how trust is lost twice. Say what is mitigated, what is still exposed, and that the underlying explanation is with the provider. Afterwards, expect your review to have a hole where their side of the timeline should be. The provider's own explanation, if you get one, usually arrives after your review is written. Write the review with the gap marked, record what you would have needed to see, and fix the part you own — because that gap is not a process failure, it is the shape of the trade you made when you chose not to operate the thing yourself.
- Your support plan promises a response within a stated period. What have you actually been promised?That someone replies within it for the severity you claimed. It says nothing about when the underlying fault is understood or fixed, and it is a separate instrument from the service's availability commitment, which is remedied with a service credit rather than with an answer.
- What do you tell stakeholders while the case sits in a queue?What is mitigated, what remains exposed, and that the explanation is with the provider with no time you control. Commit to your own mitigation and its timing; never invent an estimate for their fix, because doing so spends your credibility on someone else's schedule.
- The incident review needs a timeline and the provider has not explained their side yet. What do you do?Publish with the gap marked, stating what evidence a tenant cannot see and what you ruled out. Fix the part you own — usually instrumentation or the missing mitigation — and append the provider's account later if it arrives. Holding the review for it delays every action that is yours.
It is the building's maintenance line: you can report the burst pipe and close your own stopcock immediately, but you cannot go into the riser yourself, and the plumber arrives on their schedule.
saying these in an interview costs you the question
- Opens the provider case only after exhausting their own investigation
- Submits a case with no timestamps, identifiers or ruled-out causes
- Treats a response-time commitment as a promise of resolution
- Expects a service credit because a support reply was slow
- Has no mitigation that works without the provider's answer
- Quotes a made-up provider fix time to stakeholders