A checkout latency alert always pages the observability platform team, because the rule lives in the dashboard folder that team maintains. Why is that routing wrong, who should receive the first page instead, and how do you keep routing following ownership as services change hands?
answer
- route to levers, not to the graph
- every hop adds minutes to the clock
- ownership metadata, not a rule-name list
- paging everyone diffuses responsibility
- measure reassignment rate
basics
~20 sThe first page belongs to the team that can change the failing service, not the team that authored the rule. Route on an ownership attribute carried by the service and sourced from a service catalogue, so routing changes automatically when ownership does.
solid answer
~50 sRouting by who wrote the rule optimises for the wrong thing. The platform team can see the graph but cannot roll back checkout, so every page costs an extra hop: they wake, diagnose that it is not theirs, and forward — minutes added to every incident, and a rotation trained to act as a switchboard. The first page should go to the team that can deploy, roll back and reconfigure the affected service. I make that mechanical: the alert carries an owner attribute derived from service metadata in a catalogue, and routing keys off that attribute rather than off a hand-maintained mapping of rule names to teams. When ownership transfers, the catalogue entry changes and routing follows. Where the failing component is genuinely ambiguous, I page the owner of the user-facing journey, because they own the mitigation decision, and let them pull others in.
go deeper
Be able to say that a page should reach the team that can change the failing service, and that being able to see the dashboard is not the same as being able to fix the problem.
Explain the cost of a wrong first recipient in concrete terms — an extra hop, minutes added to the incident, a rotation trained to forward — and describe routing driven by a service ownership attribute.
Show that you keep ownership metadata authoritative by tying it to something already maintained, and that you measure reassignment rate and time to first correct responder rather than acknowledgement time.
Own the model: where ownership of record lives across the estate, how service handovers update paging, and why a central dispatch rotation is a structural mistake that keeps bad alerts alive.
## Why authorship-based routing happens, and what it costs It happens because it is the path of least resistance. A rule gets created in whatever folder or project the person creating it had open, and the destination defaults to that container's team. It looks like an administrative detail. It is not: it decides who is woken. The cost is measurable in every incident. The platform team is paged, spends several minutes establishing that the problem is inside a service they do not deploy, then forwards to the real owner, who is now starting cold with the clock already running. Two rotations lost sleep and one of them could never have fixed anything. Repeat that weekly and the platform rotation learns that its pages are usually someone else's, which is precisely the fatigue dynamic that makes them slow on the pages that *are* theirs. ## Who should get the first page The correct target is the team with the levers: it can deploy the service, roll it back, change its configuration and shed or redirect its traffic. If a team cannot do those things, paging them buys diagnosis at best and delay at worst. Two cases need care. **The failing component is not obvious.** A user-facing symptom often has several candidate causes across services. Do not page all of them. Page the owner of the user-facing journey — the service on the entry path whose users are visibly affected. That team owns the mitigation decision even when the fault turns out to be downstream, and they are the right people to pull in a dependency's owner once diagnosis narrows it. **The fault really is in shared infrastructure.** Then the platform team is the correct first recipient — but for their *own* alerts, on their own indicators, not as a receiving desk for every rule stored in their folder. The distinction is between owning the failing system and owning the tool the rule happens to live in. ## Make routing a property of the service, not of the rule The durable fix is to stop routing on rule identity at all. The alert carries an attribute naming the owning service; a catalogue maps service to owning team and its rotation; routing resolves the destination through that mapping at delivery time. The benefit is that ownership changes in exactly one place. Teams reorganise constantly — services are handed over, teams split, a squad is dissolved — and a hand-maintained list of rule-name-to-team mappings drifts silently, with the drift only discovered during the incident it makes worse. A catalogue lookup makes the transfer a data change. This is also where the fence sits against the alerting tool: the *syntax* of the routing configuration is a tool question. The practice question is what the routing keys off, and the answer is service ownership metadata that the owning team maintains as part of owning the service. ## Failure modes to name - **Page everybody.** Fanning a page to three candidate teams feels safe and produces diffusion of responsibility: each assumes another has it, and nobody acknowledges promptly. It also triples pager load for one event. - **The catch-all rotation.** A central operations rotation that receives everything and dispatches is a switchboard with a pager. It adds a hop to every incident and separates the people who feel the pain from the people who can fix the alert, which is what keeps bad alerts alive for years. - **Ownership metadata nobody maintains.** A catalogue that is not authoritative is worse than none, because routing now silently points at a team that no longer exists. Tie it to something the organisation already keeps correct — the deployment pipeline's ownership record, the repository's owners file — rather than a separate spreadsheet. - **Routing forgotten during a handover.** A service transfer that updates the code owners but not the alert routing leaves the previous team on the pager, often for months. Put routing on the transfer checklist explicitly. ## Measuring whether routing works Two numbers make this concrete and are worth quoting. First, the **reassignment rate**: what fraction of pages were forwarded to a different team after being acknowledged. A healthy service is in the low single digits; a rotation reassigning a quarter of its pages is not on-call, it is a triage desk. Second, **time to first correct responder** rather than time to first acknowledgement. Acknowledgement time flatters a system with a forwarding hop, because the wrong person acknowledged quickly. What matters is when someone who could act arrived. When a page does get reassigned, capture the reason, because it is almost always one of a small number of causes — stale metadata, a rule scoped to the wrong service, or a genuinely ambiguous symptom that needs a designated journey owner. Each has a different fix, and the reassignment log is the cheapest place to find out which one you have. ## Escalation is a separate concern One clarification worth making explicitly in an interview: getting the first recipient right is not the same as designing what happens when they do not answer. Unacknowledged-page fallbacks, secondary rotations and manager escalation are their own discipline. Routing decides who *should* have it; escalation decides what happens when they do not pick up.
- A user-facing symptom could originate in any of four services. Who gets the page?The owner of the user-facing journey — the service on the entry path whose users are visibly affected. They own the mitigation decision even if the fault turns out to be downstream, and they can pull in a dependency's owner once diagnosis narrows it. Paging all four candidates diffuses responsibility, triples pager load and typically slows acknowledgement rather than speeding it.
- What is wrong with a central operations rotation that receives all pages and dispatches them?It adds a hop to every incident and, more damagingly, separates the people who suffer bad alerts from the people who can fix them. A dispatcher cannot tune a rule they do not own, so noisy alerts survive indefinitely. Central rotations make sense for coordination roles during large incidents, not as the default first recipient for service-level pages.
- How do you notice that alert routing has drifted after a reorganisation?Track the reassignment rate — the fraction of pages forwarded to another team after acknowledgement — and log the reason each time. Drift shows up as a cluster of reassignments from one team, usually pointing at stale ownership metadata. Also put routing on the service-transfer checklist, because handovers that update code owners but not paging are the most common source.
saying these in an interview costs you the question
- Paging whoever wrote the alert rule
- Fanning one page to every candidate team
- Maintaining rule-to-team routing lists by hand
- Treating time-to-acknowledge as time-to-correct-responder
- Leaving alert routing out of a service handover