skip to content

Escalation Policies & Paging Chains

What happens when a page goes unacknowledged — automatic escalation tiers, ownership-based routing, and fallbacks. Interviewers expect you to know these concepts even if the buttons live in PagerDuty or Opsgenie.

on this pageshow

questions

5

In an on-call paging tool, what does acknowledging a page actually commit you to, and what happens if you acknowledge it and then do nothing?

level: juniorimportance: must knowfreq 50%

answer

  1. not a read receipt
  2. it stops something from happening
  3. the secondary never hears about it
  4. worse than not answering at all
  5. ownership, not acknowledgement of noise

basics

~20 s

Acknowledging a page stops the escalation chain and claims ownership: it tells the system a human is now working the problem. Acknowledging and then going back to sleep is worse than never answering, because nobody else will be notified.

solid answer

~50 s

Acknowledging is not "I saw it" — it is "I own this now." Mechanically, the ack halts the escalation timer so the secondary and the manager rungs are never notified. That is exactly why an ack followed by inaction is the most dangerous on-call behaviour there is: you have silenced the only mechanism that would have found someone else, and the outage now runs unattended with the paging system reporting it as handled. If you can't work it, don't ack — or ack and immediately escalate manually to the next person, which is a deliberate handoff rather than a silent black hole. Ack is also distinct from resolve: resolving says the condition is over, and resolving a page you have not actually fixed means the next occurrence looks like a fresh incident instead of a recurring one. Many tools guard the ack-and-sleep case with an acknowledgement timeout that re-triggers the incident if it is still unresolved after a set period.

go deeper

for a junior

Say plainly that acknowledging stops the escalation to the secondary and claims ownership, and that you never ack a page you are not about to work. Know that ack and resolve are different states.

for a middle

Explain the mechanics — the ack halts the rung timer — and describe the acknowledgement-timeout safety net that re-triggers an acknowledged but unresolved incident, including why it must be tuned so it doesn't interrupt real work.

for a senior

Talk about how you detect ack-and-sleep on a team: time-to-acknowledge paired with time-to-first-action, postmortem timelines showing a long gap after the ack, and making explicit handoffs the norm rather than a favour.

for a principal

Frame ack behaviour as a symptom of pager health. Chronic reflex-acking usually means the rotation is overloaded or the alerts are untrustworthy, and the fix is upstream reliability and alert pruning, not a policy telling people to try harder.

## Ack is a claim of ownership Every paging tool gives an incident three states that matter: triggered, acknowledged, resolved. Engineers new to on-call read "acknowledged" as a read receipt — the equivalent of marking an email seen. It is not. In an escalation policy, the acknowledgement is the signal that stops the machine from looking for anyone else. The correct mental model is: *I am awake, I have this, stop waking people.* ## What the ack does mechanically When an alert routes to a service, the escalation policy starts a timer on rung one. If the target acknowledges before the timeout, the chain halts — the secondary is never notified, the manager rung is never reached. If nobody acknowledges, the timeout expires and the page advances. So the ack is the off switch for redundancy. Pressing it is a promise that the redundancy is no longer needed, because a competent human is now on the problem. ## The ack-and-sleep failure The classic incident postmortem line is: "the page was acknowledged at 02:14; the first human action in the logs is at 03:40." What happened in between is almost always the same story — someone half-woke, hit acknowledge on their phone to make the noise stop, and fell back asleep. The result is strictly worse than not answering at all: - Had they not acked, the chain would have escalated to the secondary within minutes and someone would be working it. - Because they acked, the paging tool's dashboards say the incident is handled, so nobody escalating from another direction sees a problem either. This is why the operational rule is blunt: **do not acknowledge a page you are not about to work.** If you are unfit — too tired, in a car, mid-way through another incident — either let it escalate, or acknowledge and immediately reassign or manually escalate to the next person. A deliberate handoff is fine. A silent one is not. ## Acknowledgement timeouts and re-triggering Because the ack-and-sleep case is so common, most paging tools offer a second safety net: an acknowledgement timeout at the service level. If an incident has been acknowledged but not resolved after a configured period, the tool un-acknowledges it and restarts the escalation chain. It is a good default for high-urgency services, with one caveat — during a long incident the responder is genuinely working and will get re-paged, so the timeout should be long enough not to interrupt real work, and responders should know to expect it rather than resolve the incident just to make it stop. ## Ack versus resolve These are different claims and confusing them corrupts your incident data: - **Acknowledge** = a human owns this now. The condition is probably still true. - **Resolve** = the condition is over. Most alerting integrations resolve automatically when the underlying signal clears, which is the healthy pattern. Manually resolving a page you have not fixed makes each recurrence look like a brand-new one-off. That destroys the recurrence count, which is the evidence you need when arguing that an alert is noisy or that a service needs real remediation work. ## What good ack behaviour looks like - Ack fast, then say something in the incident channel — even "looking, nothing yet" — so humans as well as the tool know it is owned. - If it turns out to be someone else's problem, escalate or reassign explicitly rather than dropping it. - If you ack and then realise you cannot continue, hand it off out loud and let the next person ack. - Treat time-to-acknowledge as a metric worth watching, but never optimise it by acking reflexively — a good median TTA with a terrible time-to-first-action is a team quietly gaming the pager. ## Why interviewers ask this It is a cheap test for whether you have actually carried a pager. Someone who has will describe ack as ownership and will volunteer the ack-and-sleep failure unprompted. Someone who has only read about on-call describes it as a notification receipt.

  • You are paged but you're driving. What do you do?
    Don't acknowledge — let it escalate, since that is exactly what the secondary rung exists for. If I can safely reach a phone, I'd tell the secondary or the channel that I can't take it, so the handoff is explicit rather than a timeout. Acking to silence the noise would suppress the only mechanism that gets anyone else involved.
  • Why is manually resolving a page you haven't fixed harmful?
    It falsifies the incident record. Resolve means the condition cleared; if you resolve a still-broken service, the next firing looks like a fresh one-off rather than the fifth recurrence this month. That recurrence count is the evidence used to justify pruning a noisy alert or funding a real fix, and manual resolves erase it.
  • Is a low median time-to-acknowledge always a good sign?
    No. It can mean a healthy responsive rotation, or it can mean people are reflexively acking on the lock screen to stop the noise. The useful pair is time-to-acknowledge alongside time-to-first-action or time-to-mitigate; a fast ack with a slow first action is the signature of a team that has learned to silence the pager rather than answer it.

saying these in an interview costs you the question

  • Acknowledging just means you have seen the alert
  • Ack the page immediately, investigate whenever you get up
  • Resolve the incident to stop the notifications repeating
  • Acknowledging and doing nothing is the same as ignoring it
  • If it matters, the tool will keep paging after an ack

context

open as a page

You are configuring the escalation policy for a new customer-facing service in a paging tool such as PagerDuty or Opsgenie. What rungs would you define, how long would you set the acknowledgement timeout, and what must the final rung do?

level: middleimportance: must knowfreq 62%

basics

~20 s

Page the service's primary on-call first, escalate to the secondary after an unacknowledged timeout of roughly 5-15 minutes, then to a manager or a named fallback team. The final rung must reach a guaranteed-reachable human and never dead-end.

open as a page

An alert fires on a shared platform component that three product teams depend on. How do you decide which team's on-call the page routes to, and what has to exist for that routing to be correct?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Route the page to the team that owns the component and can act on it, not to everyone who depends on it. That requires an ownership record — one accountable team per service, kept next to the code and used to generate routing — plus a catch-all for alerts that map to no owner.

open as a page

Your organisation has grown to about 40 services with independent teams. Would you introduce a central first-line on-call tier that triages pages and escalates to service teams, or keep paging each service's own team directly? Argue the tradeoff.

level: principalimportance: should knowfreq 38%

basics

~20 s

Default to paging service teams directly, because the owners fix fastest and feel their own reliability. A central first-line tier buys smaller teams coverage and absorbs cross-cutting pages, but it adds a triage hop and separates the pain from the people who can remove it.

open as a page

During an incident you need help from another team that owns a dependency. When is it right to page their on-call directly instead of messaging their team channel, and what should that page contain?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

Page another team's on-call when customer impact is ongoing and you need action sooner than a channel message can deliver. The page should name the impact, the specific ask, and what you have already ruled out — enough for them to act without reading backwards.

open as a page