skip to content

How would you structure a primary/secondary on-call rotation, and what does adding a secondary actually cost the team?

level: seniorimportance: should knowfreq 38%

answer

  1. a backstop with a real job description
  2. offset the secondary by one shift
  3. two slots per cycle doubles the commitment
  4. watch the secondary's page rate
  5. shadow is not the same as secondary

basics

~20 s

A secondary is a backstop with a real job: unacknowledged pages, a second pair of hands in a declared incident, and non-urgent requests. Staffing it from the same pool doubles each person's on-call commitment, so the rotation must be deep enough.

solid answer

~50 s

My default is to staff the secondary from the same pool, offset by one shift, so this week's secondary is next week's primary — the secondary arrives at their primary shift already holding context, which makes the handoff much cheaper. The cost has to be stated plainly: with one pool covering both slots, each person is committed twice per cycle, so an eight-person rotation means one primary week and one secondary week every eight, around thirteen weeks of commitment a year. If the team cannot absorb that, the alternatives are a smaller expert or manager backstop pool, or sharing the secondary slot with an adjacent team. I also insist the secondary have a defined job — unacked pages, second responder during declared incidents, and absorbing non-urgent requests so the primary stays free — because a secondary with no job description drifts into being a second primary and you have doubled the pager load for nothing.

go deeper

for a junior

Know that a rotation usually has a primary who is paged first and a secondary who backs them up, and that both slots are named people on a published schedule, not an informal understanding.

for a middle

Be able to describe what the secondary owns — unacknowledged pages, second responder in a declared incident, non-urgent requests — and compute the commitment when both slots come from the same pool.

for a senior

Show the staffing judgment: the offset arrangement and why it helps handoffs, when to use a separate backstop pool instead, and how to read the secondary's page rate as a symptom of a broken primary rotation or over-aggressive escalation timing.

for a principal

Own the question of whether a secondary is affordable at all for a given team size, what risk you are accepting if you skip it, and how secondary duty is shared across teams for a platform without concentrating knowledge in a handful of people.

## What the secondary slot is for A secondary exists so that a page never falls on the floor and so a primary who is deep in an incident is not simultaneously the person answering everything else. In practice a well-defined secondary owns three things: pages the primary did not acknowledge; being the second responder when an incident is declared, so the primary can debug while someone else handles coordination or data-gathering; and absorbing non-urgent requests and tickets during the shift so the primary's attention stays on production. How the page actually reaches the secondary — acknowledgement timeouts, escalation tiers, fallbacks — is escalation policy design and a separate subject. What belongs to rotation design is who sits in that slot, how often, and at what cost. ## Three ways to staff it **Same pool, offset by one shift.** The most common arrangement: this week's secondary becomes next week's primary. It has a real side benefit — the incoming primary spent the previous shift adjacent to every incident, so they arrive with context and the handoff is shorter. It is also the most expensive in commitment terms. **A separate expert or manager backstop pool.** A small group of senior engineers or the manager sits behind every primary. This keeps the main rotation's commitment at one slot per cycle, but concentrates load on a few people and can entrench single points of knowledge. It works best when the secondary is genuinely rare — a true last resort rather than a routine second responder. **A shared secondary across teams.** For a platform with several small owning teams, one secondary slot can back several primaries. Cheaper per team, but the secondary needs breadth, so the runbooks have to carry the load. ## The arithmetic you must say out loud If one pool of N people covers both slots on week-long shifts, each person is on call `2 x 52 / N` weeks a year: - N = 8 → about 13 weeks a year, one primary and one secondary week every eight - N = 6 → about 17 weeks a year, roughly a third of the year touching the pager That is why a secondary is not a free safety improvement. In a six-person team, adding a secondary from the same pool is a bigger change to people's lives than most of the reliability work it is meant to protect. Either the rotation is deep enough, or the secondary comes from somewhere else, or you do without one and rely on a broader escalation path. ## Keeping it from becoming a second primary A secondary that is paged for most incidents is not a backstop; it is a second person on the same shift with half the authority. The signals are easy to read: the secondary's page rate approaches the primary's, or the secondary is routinely the one who resolves things. When that happens the cause is usually one of two, and they need different fixes. Either the primary rotation is not functioning — someone is not acknowledging, or is not equipped to act — which is a training or staffing problem. Or the escalation timing is so aggressive that the secondary is effectively paged in parallel, which is a policy problem. A rough expectation worth stating: the secondary should be paged for a small fraction of what the primary handles, and if it is not, something upstream is broken. ## Do not confuse secondary with shadow A shadow is a third arrangement entirely: someone learning the rotation who receives pages for exposure but carries no responsibility for resolving them, with a real primary always on the hook. Shadowing is a training mechanism and it does not add coverage. Counting a shadow as a secondary is how teams convince themselves they have depth they do not have. ## When to skip the secondary If the rotation is thin, incidents are rare, and the escalation path can reach a wider group quickly, a formal secondary may cost more than it returns. The honest version of that decision names the risk being accepted — an unacknowledged page takes longer to reach a human — and pairs it with something concrete, such as a shorter escalation timeout and a documented group fallback.

  • Your team is six people. Can you staff both primary and secondary from it?
    Only at about seventeen weeks of on-call commitment per person per year, which is a third of the year. I would rather share the secondary slot with an adjacent team, put a small expert or manager pool behind the primary, or drop the formal secondary and shorten the escalation timeout to a broader group. Doubling the burden on a six-person rotation buys safety by spending the thing that is already scarce.
  • What is the advantage of making this shift's secondary the next shift's primary?
    Context. The incoming primary has spent the previous shift adjacent to every incident, so they already know what is open, what is silenced and what is mid-rollout. The handoff becomes a confirmation rather than a transfer, and the first hours of the new shift are much less fragile.
  • How would you tell whether the secondary slot is doing its job?
    By the page rate ratio and what the secondary actually does. A healthy secondary is paged for a small fraction of what the primary handles, mostly unacknowledged pages and declared incidents, and spends the rest of the shift on non-urgent work. If it is being paged for most incidents, either the primary rotation or the escalation timing is broken.

saying these in an interview costs you the question

  • A secondary is free extra safety
  • Page both primary and secondary at once for speed
  • The secondary is just the person who gets woken next
  • A trainee shadowing counts as the secondary
  • Every rotation needs a secondary regardless of size

context