skip to content

You are setting up 24/7 on-call for a team that owns a production service. How many engineers does a sustainable rotation need, and what are your options if the team is smaller than that?

level: middleimportance: must knowfreq 65%

answer

  1. start from weeks per person per year
  2. 52 divided by rotation size
  3. roughly one week in six to eight
  4. leave and ramp-up shrink the real pool
  5. shrink what pages, not the humans

basics

~20 s

A single-site 24/7 primary rotation needs roughly six to eight engineers, so each person carries the pager about one week in six. With fewer people, narrow what pages overnight, merge the pager with another team, or add a second site.

solid answer

~50 s

I start with the arithmetic: with week-long primary shifts and N engineers, each person carries the pager 52/N weeks a year. Six people is about one week in six, roughly nine weeks a year, which is livable; four people is one week in four and collapses the first time someone takes leave. Google's SRE book puts the rule of thumb at around eight engineers at a single site, or six at each of two sites for a follow-the-sun team. If the team is smaller than that, I would not just push harder. I would shrink what pages overnight to symptoms that actually threaten the SLO and queue the rest as tickets, merge the pager with an adjacent team's, or run business-hours paging where the reliability target allows it. And new hires do not count until they have shadowed shifts and can act on a page alone.

go deeper

for a junior

Know that on-call is a schedule with named owners for every hour, and be able to compute your own load: 52 divided by the number of people in a weekly rotation is how many weeks a year you carry the pager.

for a middle

Be ready to size a rotation out loud, including what leave and ramp-up do to the real pool, and to name the published rule of thumb of roughly six to eight engineers per site without treating it as a law.

for a senior

Show the trade you would actually make when the team is short — narrower overnight paging tied to the reliability target, a merged pager, or a bounded period of higher load with an exit date — and explain what each choice costs.

for a principal

Own the structural question: whether the service warrants 24/7 paging at all, how many rotations the organization can genuinely staff, and when the answer is to consolidate pagers or fund a second site rather than stretch an exhausted team.

## Start from the arithmetic, not the org chart Round-the-clock coverage is 168 hours a week that must belong to a named human. If the team runs week-long primary shifts and has N people in the rotation, each person carries the pager `52 / N` weeks per year: - N = 8 → about 6.5 weeks a year, one week in eight - N = 6 → about 8.7 weeks a year, one week in six - N = 4 → 13 weeks a year, one week in four - N = 2 → 26 weeks a year, every other week, forever Then subtract reality. Vacation, sick leave, parental leave, conference travel and a new hire who is not yet cleared for the pager all come out of that pool. A six-person rotation with one person on a month of leave is a five-person rotation that month, and if you also staff a secondary from the same pool, every number above doubles. ## The published rule of thumb Google's SRE book suggests at least eight engineers at a single site for a 24/7 rotation, or six engineers at each of two sites for a team split across timezones. It is a heuristic, not physics: it is roughly the point at which per-person load stays sustainable while each person is still on call often enough to stay competent. ## Both edges fail, not just the small one **Too small.** Shifts come around too fast, there is no recovery time between them, and there is no slack to absorb a bad week or a departure. One person's holiday breaks the schedule. Worse, an on-call week is not free engineering time: at one-in-four, a quarter of the team's capacity is permanently committed to interrupts before you count the follow-up work each shift generates. **Too large.** If someone carries the pager once every twelve or more weeks, they arrive rusty — the runbooks moved, the alerts changed, the tooling was upgraded. The fix is not a smaller rotation but deliberate practice: shadow shifts, exercises, and keeping runbooks current. Very large teams usually split the rotation by service instead of leaving one pager covering everything. ## Options when you are genuinely short 1. **Reduce what pages overnight.** Keep only symptoms that threaten the reliability target as pages; everything else becomes a ticket the next shift picks up. The cost is real and must be stated: a rare overnight problem sits longer, so this is only honest if the service's target tolerates that recovery latency. 2. **Merge rotations.** One pager covering several small services owned by the same team, or a shared platform rotation across teams. The cost is breadth — the person on call now needs to act on systems they did not build, which means the runbooks have to be genuinely good. 3. **Borrow coverage.** A partner team in another timezone takes the hours you cannot staff. That is a real commitment, not a favour, and it needs shared runbooks and shared ownership. 4. **Accept a small rotation deliberately and temporarily.** With compensation, an explicit end date, and a hiring or scope plan attached. What you should not do is run a one-in-three rotation indefinitely and describe it as culture; that is how teams lose the people who know the system. ## Adding people is not instant Headcount changes the arithmetic only after the new person can be trusted alone at 3am. Budget a few shadow shifts where they receive every page with no responsibility, then a reverse-shadow where they drive and a veteran watches, and gate entry on being able to execute the top runbooks and reach every escalation path. Someone in the rotation who has to wake a colleague for every page has made the schedule look better without making the nights better. ## What an interviewer is listening for Numbers and a decision. Weeks per person per year, what leave does to the schedule, and an explicit trade you are willing to make — narrower paging scope, merged pagers, or a bounded period of higher load — rather than a general statement that on-call should be sustainable.

  • Your team is six people, but two are new hires who have never carried the pager. What is your effective rotation size, and what do you do about it?
    Effectively four, which is one week in four and not sustainable. I would keep the two new hires out of the primary slot while running them through shadow shifts and a reverse-shadow, narrow overnight paging in the meantime, and set an explicit date by which they join. Pretending they are already in the rotation just moves the failure to the first night they cannot act.
  • How do you decide whether a service deserves 24/7 paging at all?
    By what the reliability target implies about overnight recovery. If the service can absorb several hours of degradation overnight and still meet its target, it does not need a night pager — it needs a ticket queue and a documented business-hours response. Paging around the clock for a service nobody uses at 3am spends the scarcest thing the team has for nothing.
  • What happens to a weekly rotation over a holiday period, and how do you plan for it?
    Coverage thins exactly when change freezes and skeleton staffing make incidents likelier to linger. I publish the holiday schedule weeks ahead, get explicit named volunteers with compensation or time back, split long shifts into shorter ones so nobody loses a whole holiday, and confirm every escalation contact actually works before the period starts.

saying these in an interview costs you the question

  • Any team size works if people are committed enough
  • Two engineers alternating weeks is a fine rotation
  • A new hire joins the rotation on day one and counts
  • Rotation size is HR's problem, not an engineering decision
  • On-call weeks are normal project weeks with a phone

context