During an incident you need help from another team that owns a dependency. When is it right to page their on-call directly instead of messaging their team channel, and what should that page contain?
answer
- a page spends someone's night
- ongoing impact versus can-it-wait
- assume they wake up with zero context
- state the ask, not just the worry
- tell them how it ended
basics
~20 sPage another team's on-call when customer impact is ongoing and you need action sooner than a channel message can deliver. The page should name the impact, the specific ask, and what you have already ruled out — enough for them to act without reading backwards.
solid answer
~50 sThe test is whether the impact is ongoing and you need a human now. If customers are affected and you need someone with access or knowledge you don't have, page their on-call — that is what the pager is for, and hesitating out of politeness costs minutes of outage. If it can wait until working hours, it is a channel message or a ticket, because paging someone for something that could have waited is spending their night on your convenience. What separates a good page from a bad one is the payload: I'd send the impact in one line, the specific ask, the incident channel or bridge to join, and the two or three things I've already excluded. A page that just says "is the database okay?" makes the woken engineer do discovery I could have done. Afterwards, close the loop — tell them what it turned out to be, and if the page was unnecessary, say so, because that is how a cross-team paging norm survives.
go deeper
Be able to say that ongoing customer impact plus a need for their specific action justifies a page, and that a page must state impact, the ask and where to join rather than just asking someone to look.
Explain both failure directions — over-paging that exhausts neighbouring teams and under-paging that adds minutes to an outage — and describe the payload that lets a woken engineer act without reading the channel backwards.
Show the follow-through: closing the loop with the team you woke, admitting a false page, and turning the reason it looked like their fault into an action item. Distinguish paging sideways for capability from escalating up your own chain.
Make the norm structural rather than social — discoverable paging targets for every team, an explicit organisational stance that paging during impact is never taken personally, and visibility into cross-team page volume as a reliability signal.
## Why etiquette is a real topic A page costs someone their sleep or their focus, and unlike most costs in engineering it is paid by a person rather than a budget line. Teams that get this wrong fail in both directions: some page reflexively for anything touching another service, exhausting their neighbours, and some are so reluctant to page that outages run for an extra forty minutes while someone politely waits for a Slack reply. Interviewers ask this because the answer reveals whether you have been on both ends of it. ## The decision: page or message A short test, applied in order: 1. **Is there ongoing user-visible impact?** If no, do not page. A channel message or a ticket is correct even if the thing is interesting or worrying. 2. **Do I need an action only they can take?** Access you lack, a rollback of their service, a config change, knowledge of a system you have never operated. If you just want a second opinion on something you can act on yourself, that is a message. 3. **Does waiting cost more than the interrupt?** Minutes of a customer-facing outage almost always cost more than one woken engineer. Once impact is real and you need them, page immediately and stop deliberating — the hesitation is itself an outage cost. The corollary matters as much: **do not use the pager as a fast Slack.** Paging because you want an answer quickly, during business hours, from someone who is on-call but not needed, is the behaviour that makes a team stop trusting cross-team pages. ## What a page must carry The person you page wakes up with no context. Everything they need should be in the notification, because a page that requires archaeology wastes the minutes you paged to save. A usable page has four parts: - **Impact, in one line.** "Checkout is failing for roughly 30% of users since 02:10." Not "we're seeing some errors." - **The specific ask.** "We need someone who can check replication lag on the payments primary" beats "can someone from data-platform take a look." A vague ask produces a vague response. - **Where to join.** The incident channel, bridge or call, so they land in the conversation rather than starting a parallel one. - **What you have already ruled out.** Two or three lines. This is the single biggest courtesy: it stops them re-running your first fifteen minutes. Adding the severity and who is coordinating helps them calibrate how hard to push and who to talk to. ## Escalating within your own chain versus paging sideways These are different moves. Escalating within your own service's chain — pulling in your secondary or your manager — is about getting more of *your* team's capacity or authority. Paging sideways is about reaching a *capability* you do not have. Do not use a sideways page to substitute for waking your own secondary, and do not let your own escalation chain become an excuse not to reach the team who actually owns the broken thing. ## After the page Two habits keep the norm healthy: - **Close the loop.** Tell the person you paged what it turned out to be, even if their system was innocent. People who are woken and never told what happened stop responding well. - **Be honest about false pages.** If it turns out you paged them for something in your own stack, say so plainly and, where the mistake was systemic — a misleading dashboard, a missing runbook step — fix the thing that made it look like theirs. That is what stops the next team making the same call. ## The organisational version At scale you want this to be normalised rather than negotiated per incident: every team has a documented, discoverable way to be paged for genuine incidents; there is a shared expectation that a cross-team page during real impact is always acceptable and never taken personally; and cross-team pages are visible enough that a team paging its neighbours constantly becomes a conversation about their reliability rather than a private irritation.
- You paged another team and it turned out the fault was in your own service. What now?Tell them directly and quickly, and thank them — a page that ends in silence teaches people not to answer. Then look at why it looked like theirs: an ambiguous dashboard, an error message that named their service, a missing step in the runbook. Fixing that is what prevents the next false page, and it is a legitimate postmortem action item.
- Is it ever right to page a team during business hours rather than message them?Yes, when there is live customer impact and you need action now — the pager is the channel with a response guarantee, and a channel message has none. What is not acceptable is using the pager because you want a faster answer to a non-urgent question; that is borrowing an SLA you were not given, and it is how cross-team paging privileges get revoked.
- How would you make cross-team paging easy to do correctly?Make every team's paging target discoverable from the service catalog, so nobody has to ask who to wake, and publish a template for what a cross-team page must contain — impact, ask, bridge, ruled out. Then state the norm out loud: paging during real impact is always fine and never personal. Ambiguity is what makes people hesitate for twenty minutes.
saying these in an interview costs you the question
- Never page another team, it's rude to wake them
- Page whenever their service is even possibly involved
- Just write "can you take a look" and wait
- Use the pager for a faster answer during business hours
- No need to follow up once your incident is resolved