skip to content

You are configuring the escalation policy for a new customer-facing service in a paging tool such as PagerDuty or Opsgenie. What rungs would you define, how long would you set the acknowledgement timeout, and what must the final rung do?

level: middleimportance: must knowfreq 62%

answer

  1. who gets woken, and after how long
  2. the chain must terminate somewhere
  3. two or three rungs, not five
  4. five minutes to reach a laptop
  5. last rung finds a body, not the bug

basics

~20 s

Page the service's primary on-call first, escalate to the secondary after an unacknowledged timeout of roughly 5-15 minutes, then to a manager or a named fallback team. The final rung must reach a guaranteed-reachable human and never dead-end.

solid answer

~50 s

I'd build a short chain: rung one is the service's primary on-call, rung two is the secondary after an ack timeout of about five minutes for a high-urgency page, rung three is the engineering manager or a named fallback team a few minutes later. Two to three rungs is usually enough — every extra rung is minutes of unmitigated outage, and a five-rung policy just means the first four people are decoration. The timeout is a real tradeoff: too short and you wake the secondary while the primary is still fumbling for their laptop; too long and a phone on silent costs you fifteen minutes of downtime. The critical property is that the policy **terminates in a human**, not in silence — most tools let you loop the whole policy a fixed number of times, and the last rung should be someone whose job it is to find a body. I'd also test it by firing a real page before the service takes traffic.

code

yaml · 16 lines
yaml
service: checkout-api
escalation_policy:
  repeat_loops: 2
  rungs:
    - target: checkout-primary-oncall
      target_type: schedule
      ack_timeout_minutes: 5
      notify: [push, sms, phone_call]
    - target: checkout-secondary-oncall
      target_type: schedule
      ack_timeout_minutes: 10
      notify: [push, phone_call]
    - target: checkout-duty-manager
      target_type: schedule
      ack_timeout_minutes: 15
      notify: [phone_call]

go deeper

for a junior

Know the shape of a chain: primary, then secondary after an unanswered timeout, then a manager or fallback. Be able to say that acknowledging stops the escalation and that the chain must end with someone reachable.

for a middle

Explain the ack timeout as a tradeoff with a cost on both sides, justify a specific number like five minutes for a high-urgency page, and describe the notification ladder inside a rung (push, SMS, phone call).

for a senior

Show you have operated this: dead-end policies, schedules with weekend gaps, rungs pointing at disbanded teams, and the habit of firing a real test page before a service takes traffic. Tie escalation rate back to whether the primary is overloaded.

for a principal

Own the policy as an org standard rather than a per-team artifact — a default template every new service inherits, an audit that no chain dead-ends, and a rule about which services are allowed a manager rung at all versus a duty-manager pool.

## What an escalation policy is An escalation policy is the ordered list of who gets notified when an alert fires, and how long the system waits at each step before giving up on that person and moving to the next. The alert itself decides *that* a human is needed; the escalation policy decides *which* human, and what happens when that human does not answer. In tools like PagerDuty or Opsgenie this is a first-class object attached to a service, so every alert routed to that service inherits the same chain. The whole design turns on one number and one guarantee: the acknowledgement timeout, and the terminal rung. ## The rungs A standard chain for a customer-facing service looks like: 1. **Primary on-call for the owning team.** The person who is holding the pager for this service right now — targeted by schedule, never by personal name, so the policy keeps working when the rotation moves. 2. **Secondary on-call**, after the ack timeout expires. The secondary exists precisely for the case where the primary is unreachable — asleep through a phone on silent, in a tunnel, or already deep in another incident. 3. **Manager or named fallback team**, after another timeout. This rung is not there to debug. It is there because someone with authority needs to notice that a customer-facing service is broken and nobody has picked it up, and then go find people — call phones, wake the team, pull in another team. Two or three rungs is normally the right depth. Every rung is minutes of unmitigated outage, so a policy with five rungs is not more robust — it is a policy where the first four targets are decorative and the real answer arrives half an hour late. ```yaml # rung targets are schedules and teams, never individual people rungs: - target: checkout-primary-oncall # schedule ack_timeout_minutes: 5 - target: checkout-secondary-oncall # schedule ack_timeout_minutes: 10 - target: checkout-engineering-manager ``` ## Choosing the acknowledgement timeout The timeout is the time you are willing to lose to a human who is not answering. For a high-urgency page on a customer-facing path, five minutes is a common choice: it is roughly how long it takes a woken engineer to get to a laptop and press acknowledge. Ten to fifteen minutes is reasonable for lower-urgency services where the cost of a delayed response is small and the cost of unnecessarily waking a second person is not. Get it wrong in either direction and it costs you. Too short, and you burn your secondary on pages the primary was already handling — that is real pager load spent on nothing, and it trains the secondary to ignore escalations. Too long, and a single phone on silent turns a five-minute mitigation into a twenty-minute outage. If you are consistently escalating past rung one, the problem usually is not the timeout; it is that the primary is overloaded or the notification never actually arrived. ## Delivery redundancy inside a rung Escalating to another person is the coarse fallback. Inside a single rung, most tools let you configure a notification ladder for the individual: push notification immediately, SMS a minute later, phone call after that. This matters because the dominant failure is not "the engineer refused" but "the notification never landed" — a push that died with the app in the background, a carrier that dropped the SMS. A phone call that overrides do-not-disturb is what actually wakes people, so any high-urgency rung should reach a call before it escalates. ## The terminal rung The single most common defect in real escalation policies is a chain that ends in nothing. Every rung times out, the tool marks the incident unacknowledged, and no further notification is ever sent — the page is now a row in a database that nobody is looking at. Guard against it two ways: configure the policy to repeat the whole chain a fixed number of times rather than stopping after one pass, and make the last rung a target that is contractually reachable — a duty manager, a follow-the-sun partner team, or a broad team-wide notification. If your final rung is a single named individual, your escalation policy has a single point of failure with a personal life. ## Verify it before you rely on it An escalation policy is a control, and an untested control is a belief. Before a new service takes production traffic, fire a real test page and watch it walk the chain — wrong phone numbers, a schedule with a gap at the weekend, a rung pointing at a team that no longer exists, and a notification rule that only sends email are all things you discover at 3am otherwise. Re-test after any reorg, because the reorg is what breaks routing.

  • Why target a schedule rather than a specific engineer in each rung?
    Because people rotate, take leave and change teams, and a policy that names individuals silently rots. Targeting the schedule means the tool resolves whoever is genuinely on-call at page time, including overrides. Named individuals also create a single point of failure — one person on a plane and that rung is dead — and they leak responsibility to whoever happened to be on the team when the policy was written.
  • How would you handle a page for a low-urgency alert with the same policy?
    I wouldn't use the same chain. Low-urgency notifications should go to a channel or queue that gets picked up in working hours, with no ack timeout and no escalation to a person's phone. Mixing urgencies in one policy is how a disk-nearing-80% notice ends up calling the manager at 4am, and it is the fastest way to teach people that escalations don't matter.
  • What signal tells you the ack timeout is set wrong?
    Watch the rate of pages that escalate past rung one. If a meaningful share of pages reach the secondary and the primary then turns out to have been already working the issue, the timeout is too tight. If post-incident timelines repeatedly show minutes lost between the alert firing and the first human action, it is too loose — or the notification ladder never reached a phone call.

saying these in an interview costs you the question

  • More rungs makes the policy safer and more thorough
  • Ending the chain when everyone times out is acceptable
  • Naming specific engineers in each escalation rung
  • Setting one long timeout so nobody gets woken twice
  • The manager rung exists to help debug the outage

context