skip to content

An alert fires on a shared platform component that three product teams depend on. How do you decide which team's on-call the page routes to, and what has to exist for that routing to be correct?

level: seniorimportance: should knowfreq 45%

answer

  1. paging is an assertion about who can act
  2. one accountable team, not a committee
  3. notify the dependents, page the owner
  4. what happens after a reorg
  5. the alert that matches no rule

basics

~20 s

Route the page to the team that owns the component and can act on it, not to everyone who depends on it. That requires an ownership record — one accountable team per service, kept next to the code and used to generate routing — plus a catch-all for alerts that map to no owner.

solid answer

~60 s

The rule I use is: one page, one owner, and the owner is whoever can actually fix the thing the alert is about. If the shared database is saturated, that pages the platform team who owns it — fanning the same page out to all three consuming teams gives you three woken engineers, three parallel investigations and diffused responsibility, and it trains dependents to ignore pages they can't act on. Consumers get told, not paged: a channel notification or an incident-channel mention, and they page their own on-call only if their own service's SLI is breached. Making that work is a data problem more than a policy problem — you need an ownership record with exactly one accountable team per service, stored beside the code and reviewed like code, feeding a per-service routing key and escalation policy. The failure I actively hunt for is the orphan: an alert whose routing points at a team that no longer exists after a reorg. I'd route unmatched alerts to a catch-all that pages a platform on-call **and** opens a ticket to assign an owner, so orphans are loud rather than silent.

code

yaml · 15 lines
yaml
services:
  - name: payments-api
    owner_team: payments
    tier: 1
    high_urgency_route: payments-oncall
    low_urgency_route: payments-ticket-queue
  - name: shared-postgres-primary
    owner_team: data-platform
    tier: 1
    high_urgency_route: data-platform-oncall
    low_urgency_route: data-platform-ticket-queue
    consumers_notified: [payments, checkout, search]
default_route:
  owner_team: platform-oncall
  action: page_and_open_ownership_ticket

go deeper

for a junior

Know that a page should go to the team that can fix the thing, and that dependent teams are informed rather than paged. Be able to say every service needs a named owning team.

for a middle

Explain the ownership record that makes routing work — one accountable team per service, stored with the code, generating the routing key and escalation policy rather than sitting in a wiki that drifts.

for a senior

Demonstrate having been burned: orphan alerts after a reorg, schedules that resolve to nobody, fan-out pages that trained a team to ignore their pager. Describe a catch-all route that pages someone and files an ownership ticket.

for a principal

Own routing as org infrastructure — a catalog every service must be in to ship, coverage as a measurable property, and the position that shared ownership is a defect. Be ready to argue where the platform-versus-product paging boundary sits and who pays for it.

## The question behind the question When an interviewer asks how a shared dependency's page gets routed, they are checking whether you understand that paging is an ownership statement. Every page asserts: *this specific human can do something about this*. Routing that violates the assertion is how teams end up with pagers they have learned to ignore. ## One page, one owner The default is that an alert routes to the on-call of the team that owns the component the alert is about, because that team is the one that can act. For a saturated shared database, that is the platform or data team who owns the database — not the three product teams whose queries are slow because of it. The tempting alternative, fanning the page out to every dependent team, is worse in three specific ways: - **Cost.** One incident becomes three or four woken engineers, and pager load is a real budget you are spending. - **Diffusion of responsibility.** When four people are paged, each assumes one of the others has it. Ownership becomes ambiguous exactly when it needs to be sharp. - **Learned helplessness.** A team repeatedly paged for something they cannot fix stops reading those pages, and the next page they ignore will be one they could have fixed. Consumers still need to *know*. The right channel for them is notification, not paging: a message into their incident or team channel, a dependency-status update, or a mention in the incident channel. A consuming team pages its own on-call when its own user-facing SLI degrades — that is their alert firing on their service, which is a separate routing decision with a separate owner. ## Ownership metadata is the real prerequisite Routing correctness is downstream of a data problem: does a machine-readable record exist saying which team owns which service? In practice that is a service catalog or ownership file with, per service, exactly one accountable team, the paging target for that team, an escalation policy, and ideally a tier or criticality. Two properties make it stay true: 1. **It lives next to the code** and changes through review, so an ownership change is part of the change that caused it. 2. **It is the source that generates routing**, rather than being documentation that drifts from what the paging tool actually does. If the catalog and the paging tool disagree, the paging tool wins at 3am and the catalog is a lie. ```yaml services: - name: shared-postgres-primary owner_team: data-platform tier: 1 routing_key: data-platform-high-urgency - name: payments-api owner_team: payments tier: 1 routing_key: payments-high-urgency default_route: owner_team: platform-oncall action: page_and_open_ownership_ticket ``` "Exactly one" matters. Shared ownership between two teams reliably degrades to no ownership, because each team's escalation assumes the other acked. If two teams truly share a component, name one as the paging owner and let them pull the other in. ## Orphan alerts The routing defect that shows up in real postmortems is the orphan: an alert whose target is a team that was reorganised out of existence, a schedule with nobody on it, or a service that was inherited and never re-tagged. The page fires into a void and is discovered hours later. Design so orphans are loud: - Give the routing rules a **catch-all** at the end that pages a platform or infrastructure on-call, and that also opens a ticket to assign a real owner. The catch-all should be slightly embarrassing to hit, so it gets fixed. - **Audit periodically** for schedules with gaps, escalation policies whose targets no longer resolve, and services in the catalog whose owning team no longer exists in the org directory. - **Re-verify routing after every reorg.** Reorgs are the leading cause of broken paging, and nobody thinks of the pager on the day the teams change. ## Routing by urgency as well as by owner Ownership picks the team; urgency picks the delivery. The same owning team usually has at least two entry points: a high-urgency route that goes to the escalation chain and rings a phone, and a low-urgency route that lands in a queue or channel for working hours. Routing decisions are two-dimensional, and getting the owner right while sending everything down the high-urgency path is how a team ends up with a healthy ownership model and a wrecked rotation. ## What good looks like in an answer Name the ownership record, insist on a single accountable team, describe telling consumers without paging them, and volunteer the orphan case and the post-reorg audit. That combination is very hard to produce without having lived through a page that went nowhere.

  • A consuming team insists on being paged whenever the shared dependency degrades. How do you handle that?
    I'd ask what they would do with the page. If the answer is "nothing but wait," it is a notification, not a page, and I'd give them a channel feed plus a dependency status update. If they genuinely have an action — failing over to a cache, shedding load, flipping a degraded mode — then the right design is an alert on their own SLI that fires when their service is actually affected, routed to their own on-call.
  • What do you do with an alert that has no owner at all?
    Page a catch-all rather than dropping it: an infrastructure or platform on-call that acts as the owner of last resort, plus an automatically created ticket to assign a permanent owner. The catch-all must be visible and mildly painful, because a silent default route becomes a permanent home for unowned alerts and nobody ever fixes the mapping.
  • How do you keep routing correct through a reorg?
    Treat paging routes as part of the migration checklist, not an afterthought. Ownership lives in a reviewed file next to the code, so a team boundary change is a pull request that updates owners and routing keys together; then I'd fire test pages against every changed service and audit for schedules that resolve to nobody. Reorgs are the single biggest source of pages that go nowhere.

saying these in an interview costs you the question

  • Page every team that depends on the component
  • Shared ownership between two teams is fine for routing
  • Route everything to a central inbox and sort it later
  • Ownership documentation in a wiki is good enough
  • An alert with no matching route can safely be dropped

context