skip to content

On-Call Practice

Structuring sustainable on-call: rotations, escalation policies, and pager health. Interviewers ask about on-call design because it reveals how a team actually balances reliability against burnout — and whether you have lived it.

on this pageshow

questions

16

In an on-call paging tool, what does acknowledging a page actually commit you to, and what happens if you acknowledge it and then do nothing?

level: juniorimportance: must knowfreq 50%

answer

  1. not a read receipt
  2. it stops something from happening
  3. the secondary never hears about it
  4. worse than not answering at all
  5. ownership, not acknowledgement of noise

basics

~20 s

Acknowledging a page stops the escalation chain and claims ownership: it tells the system a human is now working the problem. Acknowledging and then going back to sleep is worse than never answering, because nobody else will be notified.

solid answer

~50 s

Acknowledging is not "I saw it" — it is "I own this now." Mechanically, the ack halts the escalation timer so the secondary and the manager rungs are never notified. That is exactly why an ack followed by inaction is the most dangerous on-call behaviour there is: you have silenced the only mechanism that would have found someone else, and the outage now runs unattended with the paging system reporting it as handled. If you can't work it, don't ack — or ack and immediately escalate manually to the next person, which is a deliberate handoff rather than a silent black hole. Ack is also distinct from resolve: resolving says the condition is over, and resolving a page you have not actually fixed means the next occurrence looks like a fresh incident instead of a recurring one. Many tools guard the ack-and-sleep case with an acknowledgement timeout that re-triggers the incident if it is still unresolved after a set period.

go deeper

for a junior

Say plainly that acknowledging stops the escalation to the secondary and claims ownership, and that you never ack a page you are not about to work. Know that ack and resolve are different states.

for a middle

Explain the mechanics — the ack halts the rung timer — and describe the acknowledgement-timeout safety net that re-triggers an acknowledged but unresolved incident, including why it must be tuned so it doesn't interrupt real work.

for a senior

Talk about how you detect ack-and-sleep on a team: time-to-acknowledge paired with time-to-first-action, postmortem timelines showing a long gap after the ack, and making explicit handoffs the norm rather than a favour.

for a principal

Frame ack behaviour as a symptom of pager health. Chronic reflex-acking usually means the rotation is overloaded or the alerts are untrustworthy, and the fix is upstream reliability and alert pruning, not a policy telling people to try harder.

## Ack is a claim of ownership Every paging tool gives an incident three states that matter: triggered, acknowledged, resolved. Engineers new to on-call read "acknowledged" as a read receipt — the equivalent of marking an email seen. It is not. In an escalation policy, the acknowledgement is the signal that stops the machine from looking for anyone else. The correct mental model is: *I am awake, I have this, stop waking people.* ## What the ack does mechanically When an alert routes to a service, the escalation policy starts a timer on rung one. If the target acknowledges before the timeout, the chain halts — the secondary is never notified, the manager rung is never reached. If nobody acknowledges, the timeout expires and the page advances. So the ack is the off switch for redundancy. Pressing it is a promise that the redundancy is no longer needed, because a competent human is now on the problem. ## The ack-and-sleep failure The classic incident postmortem line is: "the page was acknowledged at 02:14; the first human action in the logs is at 03:40." What happened in between is almost always the same story — someone half-woke, hit acknowledge on their phone to make the noise stop, and fell back asleep. The result is strictly worse than not answering at all: - Had they not acked, the chain would have escalated to the secondary within minutes and someone would be working it. - Because they acked, the paging tool's dashboards say the incident is handled, so nobody escalating from another direction sees a problem either. This is why the operational rule is blunt: **do not acknowledge a page you are not about to work.** If you are unfit — too tired, in a car, mid-way through another incident — either let it escalate, or acknowledge and immediately reassign or manually escalate to the next person. A deliberate handoff is fine. A silent one is not. ## Acknowledgement timeouts and re-triggering Because the ack-and-sleep case is so common, most paging tools offer a second safety net: an acknowledgement timeout at the service level. If an incident has been acknowledged but not resolved after a configured period, the tool un-acknowledges it and restarts the escalation chain. It is a good default for high-urgency services, with one caveat — during a long incident the responder is genuinely working and will get re-paged, so the timeout should be long enough not to interrupt real work, and responders should know to expect it rather than resolve the incident just to make it stop. ## Ack versus resolve These are different claims and confusing them corrupts your incident data: - **Acknowledge** = a human owns this now. The condition is probably still true. - **Resolve** = the condition is over. Most alerting integrations resolve automatically when the underlying signal clears, which is the healthy pattern. Manually resolving a page you have not fixed makes each recurrence look like a brand-new one-off. That destroys the recurrence count, which is the evidence you need when arguing that an alert is noisy or that a service needs real remediation work. ## What good ack behaviour looks like - Ack fast, then say something in the incident channel — even "looking, nothing yet" — so humans as well as the tool know it is owned. - If it turns out to be someone else's problem, escalate or reassign explicitly rather than dropping it. - If you ack and then realise you cannot continue, hand it off out loud and let the next person ack. - Treat time-to-acknowledge as a metric worth watching, but never optimise it by acking reflexively — a good median TTA with a terrible time-to-first-action is a team quietly gaming the pager. ## Why interviewers ask this It is a cheap test for whether you have actually carried a pager. Someone who has will describe ack as ownership and will volunteer the ack-and-sleep failure unprompted. Someone who has only read about on-call describes it as a notification receipt.

  • You are paged but you're driving. What do you do?
    Don't acknowledge — let it escalate, since that is exactly what the secondary rung exists for. If I can safely reach a phone, I'd tell the secondary or the channel that I can't take it, so the handoff is explicit rather than a timeout. Acking to silence the noise would suppress the only mechanism that gets anyone else involved.
  • Why is manually resolving a page you haven't fixed harmful?
    It falsifies the incident record. Resolve means the condition cleared; if you resolve a still-broken service, the next firing looks like a fresh one-off rather than the fifth recurrence this month. That recurrence count is the evidence used to justify pruning a noisy alert or funding a real fix, and manual resolves erase it.
  • Is a low median time-to-acknowledge always a good sign?
    No. It can mean a healthy responsive rotation, or it can mean people are reflexively acking on the lock screen to stop the noise. The useful pair is time-to-acknowledge alongside time-to-first-action or time-to-mitigate; a fast ack with a slow first action is the signature of a team that has learned to silence the pager rather than answer it.

saying these in an interview costs you the question

  • Acknowledging just means you have seen the alert
  • Ack the page immediately, investigate whenever you get up
  • Resolve the incident to stop the notifications repeating
  • Acknowledging and doing nothing is the same as ignoring it
  • If it matters, the tool will keep paging after an ack

context

open as a page

You are configuring the escalation policy for a new customer-facing service in a paging tool such as PagerDuty or Opsgenie. What rungs would you define, how long would you set the acknowledgement timeout, and what must the final rung do?

level: middleimportance: must knowfreq 62%

basics

~20 s

Page the service's primary on-call first, escalate to the secondary after an unacknowledged timeout of roughly 5-15 minutes, then to a manager or a named fallback team. The final rung must reach a guaranteed-reachable human and never dead-end.

open as a page

How do you measure whether an on-call rotation is healthy, and what pager-load numbers would tell you it is not?

level: middleimportance: must knowfreq 65%

basics

~20 s

Measure pager load per shift, not per month: pages per shift, the share that arrive outside business hours, the share that turned out to be actionable, and hours lost to interrupts. Google's SRE book suggests at most two incidents per 12-hour shift.

open as a page

You are setting up 24/7 on-call for a team that owns a production service. How many engineers does a sustainable rotation need, and what are your options if the team is smaller than that?

level: middleimportance: must knowfreq 65%

basics

~20 s

A single-site 24/7 primary rotation needs roughly six to eight engineers, so each person carries the pager about one week in six. With fewer people, narrow what pages overnight, merge the pager with another team, or add a second site.

open as a page

What has to transfer at an on-call shift handoff, and what goes wrong when the handoff is just a message saying "quiet week"?

level: middleimportance: must knowfreq 58%

basics

~20 s

A handoff transfers state, not a summary: open incidents and their next step, every alert silenced during the shift with its expiry, in-flight changes and migrations, temporary manual fixes with an owner, and risky events scheduled in the next shift.

open as a page

Your team's on-call rotation is taking roughly 50 pages a week and two engineers have asked to come off it. What do you do in the first two weeks, and what do you change structurally?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Start with data, not sympathy: rank every alert by how many pages it produced last month, then give each of the top offenders a disposition — delete, demote to a ticket, or fund the reliability fix — with a named owner and a date. Structurally, agree a pager-load cap with a pre-committed consequence.

open as a page

Would you run on-call as week-long shifts or as daily or 12-hour shifts, and what drives that choice?

level: middleimportance: should knowfreq 45%

basics

~20 s

Week-long shifts give context continuity and one handoff per week, but concentrate every bad night on one person. Twelve-hour or daily shifts contain fatigue and let awake people take night pages, at the cost of many more handoffs.

open as a page

An alert fires on a shared platform component that three product teams depend on. How do you decide which team's on-call the page routes to, and what has to exist for that routing to be correct?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Route the page to the team that owns the component and can act on it, not to everyone who depends on it. That requires an ownership record — one accountable team per service, kept next to the code and used to generate routing — plus a catch-all for alerts that map to no owner.

open as a page

Why do experienced teams track on-call interrupt hours rather than page counts alone, and what does a single overnight page actually cost the next day?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Page counts treat a two-minute acknowledgement and a four-hour incident as the same event. Interrupt hours capture triage, recovery and the lost re-focus time around each page, and an overnight page costs the incident plus a degraded next day, which is why some conditions should be tickets rather than pages.

open as a page

How would you structure a primary/secondary on-call rotation, and what does adding a secondary actually cost the team?

level: seniorimportance: should knowfreq 38%

basics

~20 s

A secondary is a backstop with a real job: unacknowledged pages, a second pair of hands in a declared incident, and non-urgent requests. Staffing it from the same pool doubles each person's on-call commitment, so the rotation must be deep enough.

open as a page

Your organisation has grown to about 40 services with independent teams. Would you introduce a central first-line on-call tier that triages pages and escalates to service teams, or keep paging each service's own team directly? Argue the tradeoff.

level: principalimportance: should knowfreq 38%

basics

~20 s

Default to paging service teams directly, because the owners fix fastest and feel their own reliability. A central first-line tier buys smaller teams coverage and absorbs cross-cutting pages, but it adds a triage hop and separates the pain from the people who can remove it.

open as a page

How would you turn an on-call pager-load cap into a policy the organization actually respects, rather than a number on a dashboard?

level: principalimportance: should knowfreq 38%

basics

~20 s

A cap becomes policy only when a breach triggers a consequence agreed in advance, with a named decider and an expiry on exceptions. Pre-commit while things are calm, publish the number routinely, and pair it with a detection metric so nobody meets the cap by silencing alerts.

open as a page

When is follow-the-sun on-call worth building instead of a single-region rotation that takes night pages?

level: principalimportance: should knowfreq 34%

basics

~20 s

Follow-the-sun is worth it when night paging is irreducible and the organization already has real engineering teams in other timezones. Otherwise it is a hiring strategy, not an on-call policy, and alert pruning is far cheaper.

open as a page

What are the common ways companies compensate engineers for carrying an on-call pager, and why do SRE teams often cap that compensation?

level: juniorimportance: nice to knowfreq 30%

basics

~20 s

Two mainstream forms exist: cash — a per-shift stipend, sometimes plus payment for incident time — and time off in lieu. Google's SRE book describes capping either at a small proportion of salary, so that a noisy pager never becomes income anyone has a reason to protect.

open as a page

In an on-call schedule, what is an override, and why is agreeing in a chat message to cover for a teammate not the same thing?

level: juniorimportance: nice to knowfreq 30%

basics

~20 s

An override replaces the on-call person for a defined window in the scheduling system, so pages actually route to the covering engineer. An informal agreement changes nothing: the paging tool still calls the original person's phone.

open as a page

During an incident you need help from another team that owns a dependency. When is it right to page their on-call directly instead of messaging their team channel, and what should that page contain?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

Page another team's on-call when customer impact is ongoing and you need action sooner than a channel message can deliver. The page should name the impact, the specific ask, and what you have already ruled out — enough for them to act without reading backwards.

open as a page