skip to content

Your organisation has grown to about 40 services with independent teams. Would you introduce a central first-line on-call tier that triages pages and escalates to service teams, or keep paging each service's own team directly? Argue the tradeoff.

level: principalimportance: should knowfreq 38%

answer

  1. who feels the pain versus who can fix it
  2. every hop costs minutes of outage
  3. coverage math for a small team
  4. cross-cutting alerts have no service owner
  5. absorbed pain is unfixed pain

basics

~20 s

Default to paging service teams directly, because the owners fix fastest and feel their own reliability. A central first-line tier buys smaller teams coverage and absorbs cross-cutting pages, but it adds a triage hop and separates the pain from the people who can remove it.

solid answer

~60 s

I'd keep direct-to-owner paging as the default. The team that wrote the service diagnoses it fastest, and — more importantly — the pain lands on the people with the power to remove it, which is the feedback loop that drives reliability work. A first-line tier breaks that loop: it inserts minutes of triage latency into every incident, and it lets a team ship unreliable code while someone else absorbs the nights. Where a central tier genuinely earns its place is narrower than people expect: covering hours no single small team can staff, owning genuinely cross-cutting signals like edge, network or capacity where no one service is the owner, and being the owner-of-last-resort for unmatched alerts. The hybrid I'd actually build is direct paging for service-specific alerts, a central tier for cross-cutting and orphan routes, and a hard rule that the central tier never absorbs a recurring service-specific page — that goes back to the owning team, because absorbed pain is unfixed pain. If a team is too small to hold a pager, that is a team-size problem, not a paging problem.

go deeper

for a junior

Know the two models by name and the basic tradeoff: the owning team fixes fastest, while a central tier provides coverage. You are not expected to argue the org-level economics.

for a middle

Be able to state both costs concretely — triage latency added to every incident versus an unsustainable rotation for a small team — and explain why cross-cutting alerts have no natural service owner.

for a senior

Argue the feedback loop: pain landing on the people who can remove it is what actually drives reliability work, and a tier that absorbs recurring pages removes the pressure to fix them. Propose the hybrid split with a rule about what the tier will not absorb.

for a principal

Own the decision with evidence and a review point — where postmortem minutes actually go, the tier's escalate-without-acting rate, per-team page trends — plus the org consequences: staffing, retention on the tier, runbook upkeep, and the policy that returns absorbed pages to owners.

## The two models **Direct paging ("you build it, you run it").** Each service's alerts route to the owning team's escalation policy. The people who wrote the code get woken by it. **Central first-line tier.** A dedicated rotation — historically an NOC, today more often a platform or production-engineering on-call — receives pages first, triages against a runbook, mitigates what it can, and escalates to the owning team when it cannot. Both exist in serious organisations. The interview is not looking for a doctrine; it is looking for whether you can price each one. ## What direct paging buys - **Time to mitigate.** The owner needs no handoff, knows what shipped yesterday, and can act on their own systems without asking permission. - **The feedback loop.** This is the real argument. When the team that ships the code holds the pager, unreliability has a personal cost to exactly the people who can remove it. Noisy alerts get pruned, flaky retries get fixed, and "we'll clean it up later" becomes "I am not doing that again next Tuesday." - **Ownership clarity.** Everyone knows who has it, always. ## What direct paging costs - **Coverage.** A rotation needs enough people to be sustainable — commonly cited as six or more so nobody carries the pager more than roughly one week in six. A four-person team paging 24/7 is a burnout schedule, and a follow-the-sun equivalent needs staff in multiple regions. - **Duplication.** Forty teams each build the same on-call muscle, and forty pagers means the same cross-cutting incident can wake many people at once. - **Uneven maturity.** Some teams will be excellent at it and some will not, and their customers experience the difference. ## What a first-line tier buys - **Coverage without staffing every team for nights.** One rotation of specialists covers the hours that dozens of small teams cannot. - **A home for cross-cutting signals.** Edge health, network, shared capacity, region-level failure — alerts where no single service is the owner and the correct first responder is someone with a whole-system view. - **An owner of last resort.** Unmatched alerts and orphaned services have somewhere to land, loudly, rather than firing into a void. - **Consistency.** One rotation with drilled process, tested runbooks and a known way to run an incident, which is genuinely better than forty ad-hoc versions. ## What it costs — and this is the part candidates miss - **A triage hop on every incident.** Minutes spent by someone reading a runbook before the person who knows the answer is even aware. For a tier-1 outage that latency is the dominant cost. - **Broken feedback.** The team shipping the code stops feeling the consequences. Alert noise nobody suffers from is alert noise nobody prunes; a service that pages twice a night becomes someone else's problem indefinitely. This is the failure that quietly kills reliability programmes. - **A runbook maintenance treadmill.** The tier can only act on what is written down, and the writing is done by teams who no longer feel the pain of it being stale. - **A career and morale hazard.** A rotation whose job is to route pages it cannot fix is corrosive work; retention there is often poor, which makes the tier's quality unstable. ## The hybrid I would actually build 1. Service-specific alerts route **direct to owning teams**. This is the default and the majority of pages. 2. A **central tier owns cross-cutting signals** — edge, network, shared infrastructure, region-wide capacity — plus the catch-all for unrouted alerts, and acts as incident-command support for large multi-team incidents. 3. A **hard rule with teeth: the central tier does not absorb recurring service-specific pages.** If a service pages the central tier repeatedly, the alert routes back to the owning team with the recurrence data attached. Absorbed pain is unfixed pain, and this rule is the whole defence of the model. 4. **Team-size problems get solved as team-size problems.** If a team is too small to hold a pager sustainably, merge rotations across two or three related teams so people are still paged for systems they know, rather than pushing the burden to strangers. ## How I would decide, and how I would know it was working Evidence beats preference. I would look at where minutes actually go in postmortem timelines: if time-to-mitigate is dominated by triage latency, the central tier is hurting; if it is dominated by nobody being awake in a region, coverage is the binding constraint. I would watch what share of pages the first-line tier resolves without escalating — a low share means it is an expensive relay — and I would watch per-team page volume trends, because a healthy model shows them falling as teams fix what wakes them. A central tier whose own page volume grows year over year is not scaling the org; it is laundering unreliability, and that is the point at which I would push routing back to owners.

  • Where does a central first-line tier clearly beat direct paging?
    Cross-cutting signals with no single service owner — edge, network, shared infrastructure, region-level failure — plus the catch-all for alerts that route nowhere, and coordination support on large multi-team incidents. Those are cases where the first responder needs a whole-system view rather than deep knowledge of one service, and where direct routing has no correct target to pick.
  • A six-person team says it cannot sustain 24/7 on-call. What is your answer?
    Treat it as a coverage problem, not a routing problem. Options in order: merge rotations with adjacent teams so responders are still paged for systems they understand; reduce the alerts that fire out of hours until only genuine impact pages; or, if the service is not truly tier-1, stop paging for it overnight and let it wait. Handing the pager to strangers is the last resort because it severs the feedback loop.
  • What metric would tell you a first-line tier has become a liability?
    The share of pages it escalates without acting. If most pages pass straight through, it is an expensive relay adding triage latency to every incident. I'd pair that with per-service page recurrence: services whose page volume is flat or rising while the owning team never feels it are being subsidised, and their routing should go back to the owner with the data attached.

saying these in an interview costs you the question

  • A central NOC tier is always the mature choice at scale
  • Direct paging doesn't scale past a handful of teams
  • Adding a triage tier has no effect on time to mitigate
  • Small teams should hand their pager to a central rotation
  • Whoever answers first matters more than who owns the service

context