Your team is adding a new alert for a production service and must decide whether it pages a human immediately, files a ticket, or is only recorded for later investigation. What test decides the tier, and what does each of the three tiers commit the team to?
answer
- three tiers, three different promises
- cost of an interrupt is real
- impact now, and a lever to pull
- same script every time means automate
- informational page is a contradiction
basics
~20 sPage only when a human must act within minutes and can actually change the outcome. Ticket when the work is real but can wait days. Log when nobody will act. Each tier is a response-time promise.
solid answer
~50 sI treat the three tiers as promises about human response, not as labels. A page promises that someone is interrupted now, acknowledges in minutes and starts working; a ticket promises an owner and action within the normal week; a log promises nothing at all except that the signal exists when someone goes looking. So the test is: is a user or the service's objective being harmed now or on a trajectory that will harm them before the next working day, and is there something a human can decide or do right now that changes the outcome? If the impact is not there, it is a ticket. If the impact is there but the action is always the same mechanical fix, the honest answer is to automate it rather than page. Anything labelled an "informational page" is a contradiction and should be demoted.
go deeper
Be able to say plainly that a page means someone is woken and expected to act within minutes, and that anything nobody would act on tonight belongs in a ticket instead.
Explain the actionability test end to end: real or imminent impact, a lever a human can pull, and a decision that is not purely mechanical. Give a concrete example of each tier.
Show that you treat over-paging as a detection failure rather than a comfort issue, and describe how you demote rules using recorded evidence that responders took no action.
Own the policy: who may promote an alert to paging tier, how the tier is reviewed like code, and how you keep a fleet-wide standard without becoming the bottleneck every team routes around.
## A tier is a promise, not a label An alerting policy is a set of promises about human behaviour, and the three tiers differ only in what they promise. - **Page.** Someone is interrupted immediately — woken at 03:00, pulled out of a meeting — and is expected to acknowledge within minutes and start working the problem. This is the expensive tier: an interrupt plus the context switch around it costs a responder the better part of an hour even when the alert turns out to be nothing, and a night page costs part of the following day as well. - **Ticket.** The work is real and gets a named owner, but it is picked up in the normal flow of the week or sprint. Nobody's evening is spent on it. - **Log or dashboard panel.** No promise at all. The signal is recorded so that a human already investigating something can find it. Nobody is expected to look at it on its own. Written down this way, tiering stops being a matter of taste. You are not asking "is this important?" — almost everything feels important to the team that instrumented it. You are asking "what am I willing to commit a human to?" ## The actionability test Three questions, in order. 1. **Is impact real or imminent?** Is a user, a customer commitment or the service's stated objective being harmed right now, or on a trajectory that will harm it before the next working day? If the answer is no, the highest tier this deserves is a ticket, no matter how alarming the graph looks. 2. **Can a human change the outcome right now?** If there is no lever — no rollback, no failover, no capacity to add, no traffic to shed, no communication to send — then paging changes nothing except who is awake. Note that "decide to declare an incident and tell customers" is a genuine lever; "watch it" is not. 3. **Does it need a human at all?** If the runbook's entire content is "run this script", the alert is not really telling you about the system, it is telling you about a missing piece of automation. Page-then-run-the-same-script is a work queue with a pager attached. Fail question 1 and it is a ticket. Fail question 2 and it is a ticket or a log line. Fail question 3 and it is an automation backlog item, and until that automation exists you have consciously chosen to keep paying a human for it. ## The cost of getting it wrong, in both directions Over-paging is the common failure and the more insidious one. A responder who is paged for things that need no action learns, correctly, that the pager is usually wrong. The rule keeps firing and the detection still "works" on paper, but the human recall collapses: the one page in ten that mattered is acknowledged and set aside with the others. Alert fatigue is not a morale problem, it is a detection problem. Under-paging is rarer but sharper. Something that genuinely needed a person sat in a queue until a customer reported it, and the incident's clock started from the customer's complaint rather than from the signal you already had. The tell in a postmortem is a timeline where the monitoring data existed the whole time and nobody was told. ## Runway makes the boundary quantitative The hardest tier calls are slow-moving conditions, and there the useful framing is runway. A resource that will exhaust in three weeks is a ticket; the same condition with twenty minutes of runway is a page, because twenty minutes is comparable to the time it takes to wake a human, get them oriented and let them act. The threshold that matters is not the level of the metric but whether the remaining time is longer or shorter than a human response cycle. State the crossover explicitly in the alert definition so it can be argued about later. ## Keeping the tiers honest over time Tiers rot, because systems change and rules do not. Two habits keep them honest. First, after every page the responder records whether they actually took an action; a rule with a persistent no-action rate is a demotion candidate and someone has to be empowered to make that call. Second, the tier lives in the alert definition next to the rule, reviewed like code, so that promoting something to paging tier is a visible decision with a reviewer rather than a default. ## A worked set Checkout error rate at 2% sustained, with real users failing to pay: page — impact is live, and rollback is a lever. A nightly batch job that failed but will retry tomorrow with no user-visible effect: ticket. A node crossing 70% disk with weeks of headroom: ticket; the same disk filling in under an hour: page. A third-party provider degraded with no failover available: page the person who owns the decision to communicate, not the whole team to watch a graph together.
- A team argues that a particular condition must page because "if we don't page we'll never look at it". How do you respond?That is an argument that their ticket queue is untrusted, not that the condition needs a human at 03:00. Paging to compensate for a broken backlog spends the on-call's sleep to fix a process problem, and it degrades every other page on the rotation. Fix the queue: give the ticket an owner and a service-level expectation, and revisit the tier if impact turns out to be real.
- Where do you draw the line between a page and a ticket for a slowly exhausting resource?On runway rather than level. If the remaining time is longer than a human response cycle — waking someone, orienting them, acting — a ticket is enough and the threshold should be set far enough out that the fix fits in a working day. Once the projected exhaustion is inside that window, it becomes a page, because nobody will be there to act otherwise.
- What does it mean if an alert's runbook consists of one command?It means the alert is a queue of automation work, not a decision point. A human adds nothing except latency and error. The right move is to run that remediation automatically and page only if the automated attempt fails or repeats abnormally often, which converts a recurring interrupt into a rare, genuinely informative one.
saying these in an interview costs you the question
- Paging anything that looks unusual on a graph
- Treating an informational page as a valid tier
- Assuming importance to the team equals paging urgency
- Paging because the ticket queue is ignored
- Keeping a page whose runbook is one command