How would you turn an on-call pager-load cap into a policy the organization actually respects, rather than a number on a dashboard?
answer
- a number without a consequence is decoration
- pre-commit before the bad quarter
- exceptions need approver and expiry
- publish a counterweight against gaming
- ratchet the cap, don't copy it
basics
~20 sA cap becomes policy only when a breach triggers a consequence agreed in advance, with a named decider and an expiry on exceptions. Pre-commit while things are calm, publish the number routinely, and pair it with a detection metric so nobody meets the cap by silencing alerts.
solid answer
~50 sA number without a consequence is a dashboard, not a policy. I would pre-agree three things before any breach: the cap itself in shift terms, what automatically happens when it is exceeded, and who is allowed to grant an exception and for how long. The consequence needs teeth and a real cost — the usual candidates are that alert cleanup and reliability work take precedence over roadmap commitments until the rotation is back under the cap, or, in a stronger model, that operational responsibility returns to the team that shipped the service until it meets the standard. Pre-commitment matters because arguing for relief in the middle of a bad quarter always loses to shipping pressure. Then guard against gaming: publish the actionable rate and customer-reported detections alongside pager load, so hitting the cap by turning alerts off is visible immediately. And put the number in a routine forum, not an escalation, so it is boring by the time it matters.
go deeper
Understand that a pager-load target only means something if breaching it changes what the team works on next, and know where your team's number is published.
Explain why a cap needs a rate over a window rather than a single-shift trigger, and how a counterweight metric stops the number being met by silencing alerts.
Show you can operate the policy: run the routine review, apply the disposition work when the cap is breached, and record exceptions with an owner and an expiry rather than letting them drift.
Own the incentive design end to end — securing the sponsor, pre-committing the consequence before a crisis, choosing an enforceable starting number for your actual staffing, and ratcheting it down as reliability work lands.
## Why caps fail Most teams that have an on-call health problem also already have a number. Someone put "max 2 pages per shift" in a wiki. It changes nothing, and the reason is structural rather than cultural: the cap describes a state, but nothing in the organisation *does* anything when the state is violated. When the pager is bad, the team is also behind on delivery, and the person who would have to enforce the cap is the person under delivery pressure. The cap loses every time. So the design problem is not choosing the number. It is arranging for the breach to have a consequence that does not depend on someone finding the courage in a bad week. ## The three pre-commitments **1. The threshold, in shift terms.** Pages per shift, expressed as a rate over a window — for example, more than a stated number of shifts in a rolling month exceeded the limit. Rate over a window beats a single-shift trigger, which fires on one bad night and gets ignored. **2. The automatic consequence.** This is the whole policy. Options, roughly in increasing strength: - *Cleanup takes precedence.* Until the rotation is back under the cap, alert cleanup and reliability fixes rank above roadmap work. The cost is explicit: named commitments slip. - *Feature launches pause for the affected service.* Stronger, because it also stops adding new failure modes to a rotation that cannot absorb them. - *Operational ownership reverts.* The service's pager routes back to the team that built it until it meets the agreed operational standard. This is the model Google's SRE organisation is known for, and it works precisely because SRE support is a service that can be withdrawn. It only exists as a lever where a separate operating team exists in the first place; be honest in an interview about whether it applies to the org you are describing. **3. The exception path.** There will be legitimate exceptions — a migration, a launch. So define them: exceptions require a named approver, a stated reason, and an expiry date. An exception without an expiry is a repeal. ## The gaming problem Any metric attached to a consequence gets optimised, including in ways you did not intend. Pager load can be improved without improving anything by silencing alerts, raising thresholds until nothing fires, or reclassifying pages as tickets that nobody reads. This is not a hypothetical; it is the default outcome if the cap stands alone. The defence is a counterweight metric published beside it, so the trade is visible: - **Actionable rate** — falling pager load with a stable-or-rising actionable rate is genuine improvement. - **Customer-reported and externally detected incidents** — if these climb while pages fall, detection was traded away. - **Reliability outcomes** — user-visible failures should not be getting worse while the pager gets quieter. A cap paired with a detection metric is hard to game; a cap alone is trivially gameable. ## Making it boring The policies that survive are the ones that appear in a routine forum. Pager load should be reported in the same review as delivery status, every cycle, whether or not it is bad. Two benefits: the trend is visible before it is a crisis, and by the time a breach occurs the number is familiar rather than something the team is suddenly waving in an escalation. Anything that only ever gets raised as a complaint reads as a complaint. It also needs a sponsor with the authority to let commitments slip. A cap enforced by an engineer against a director's roadmap is not a policy; it is an argument that engineer will lose. ## Choosing the number Copying an external figure uncritically is a weak answer. Google's SRE guidance of at most two incidents per 12-hour shift comes with a staffing model behind it. A team of six covering everything themselves may have to start from where they are: set the cap slightly below the current p90 shift, enforce it, and ratchet it down each quarter. A cap you can actually hold and lower beats an aspirational number that is breached from day one and therefore ignored. ## What the interviewer is checking That you understand incentives, not that you can recite a threshold. The strong answer has: a pre-agreed consequence with a real cost, a named decider, expiring exceptions, a counterweight metric against gaming, and a sponsor. The weak answer is a number and a hope that people will care.
- What is the strongest consequence you have seen attached to a pager-load breach, and when does it not apply?Reverting operational ownership — the service's pager goes back to the team that built it until it meets the agreed standard. It works because on-call support is framed as a service that can be withdrawn. It does not apply where the same team builds and operates everything, which is most organisations; there the realistic lever is that cleanup outranks roadmap commitments until the rotation recovers.
- A team hits the cap for the first time during a critical launch quarter. Do you enforce it?Use the exception path rather than either ignoring the cap or blocking the launch. A named approver grants a time-boxed exception with a stated end date and an agreed cleanup commitment for after the launch. Silently not enforcing teaches everyone the cap is decorative; a recorded, expiring exception keeps the mechanism intact and makes the debt visible.
- How would you set the cap for a six-person team that cannot copy Google's staffing model?Start empirically. Measure the current distribution and set the cap a little below today's p90 shift, so it is breachable but not breached constantly, then ratchet it down each quarter as cleanup lands. An enforceable number that improves beats an aspirational one that is violated from the first week and quietly ignored thereafter.
saying these in an interview costs you the question
- Setting a cap with no defined consequence for breaching it
- Allowing open-ended exceptions with no approver or expiry
- Publishing pager load without any anti-gaming counterweight
- Copying Google's numbers without their staffing assumptions
- Raising the cap only during escalations, never in routine reviews