skip to content

When is follow-the-sun on-call worth building instead of a single-region rotation that takes night pages?

level: principalimportance: should knowfreq 34%

answer

  1. daylight owns the pager
  2. two sites twelve hours apart
  3. each site needs a full rotation
  4. the token outpost failure
  5. prune the nights before buying a continent

basics

~20 s

Follow-the-sun is worth it when night paging is irreducible and the organization already has real engineering teams in other timezones. Otherwise it is a hiring strategy, not an on-call policy, and alert pruning is far cheaper.

solid answer

~50 s

Follow-the-sun means each site is on call only during its own business day, so nobody is paged at 3am. That is a genuine step change in on-call quality, but the price is high: you need two sites roughly twelve hours apart, or three at eight, each with enough engineers to run its own rotation — the SRE book's rule of thumb is about six per site — plus duplicated ownership knowledge, shared runbooks and two or three handoffs every day. The classic failure is a token two-person outpost that cannot actually resolve anything and escalates back to headquarters overnight: you have paid for a second site and kept the night pages. So the decision hinges on two things. How much irreducible night paging is there really, and does the company already have engineering presence in another region? If nights are quiet or the noise is fixable, pruning alerts and improving reliability is dramatically cheaper than standing up a continent.

go deeper

for a junior

Know what the model means in practice — each site holds the pager only during its own working day — and that it exists to remove night pages rather than to spread work around.

for a middle

Be able to state the mechanics and their costs: roughly six engineers per site, two or three handoffs a day, and runbooks and access that must work identically at every site.

for a senior

Show that you would count the actual out-of-hours pages first and separate noise and fixable reliability issues from irreducible events, because the cheaper levers usually cover most of the volume.

for a principal

Own the framing that this is an org-design and hiring commitment, not an on-call policy: argue it on the irreducible remainder, insist the second site be a peer with real ownership, and be ready to say no when the honest answer is to fix the alerting.

## What follow-the-sun actually is In a follow-the-sun model, on-call responsibility moves around the globe with daylight. A site in one region holds the pager during its working day and hands it to the next region at the end of that day. Each engineer is only ever on call during hours they would be awake and working anyway. Nobody is woken at 3am, because at 3am the pager belongs to a team having breakfast. ## What it costs **Real teams, not outposts.** Each site needs enough engineers to run a full rotation of its own — roughly six per site is the published rule of thumb for a two-site arrangement, versus about eight for a single site covering everything. That means hiring, retaining and growing an actual engineering team in another region, with genuine ownership of the service. **Duplicated knowledge.** Every site must be able to resolve the incidents that arrive during its window. That means runbooks that work for someone who did not build the system, access and permissions granted everywhere, and deliberate investment in keeping the second site's expertise current. Knowledge decays fastest at the site that sees the fewest incidents. **Handoff seams.** Two sites means two handoffs a day; three sites means three. Over a year that is hundreds of moments where an open silence, an in-flight change or an unfinished investigation can be dropped. Incidents that span a handoff need the response role transferred explicitly rather than evaporating at the end of a working day. **Organizational overhead.** Aligned tooling, a shared incident vocabulary, comms that work across languages and holidays, and a decision structure that does not force every non-trivial call back to one headquarters timezone. ## The characteristic failure The most common way this model fails is the token site: two or three engineers hired in another region, given the pager, but never given ownership, depth or the authority to change anything. Every real incident escalates back to the original team — at 3am local time, exactly as before, now with an extra hop and a language barrier in the middle. The organization has paid for a second site and bought no nights back. If you cannot staff a site to the point where it resolves incidents autonomously, follow-the-sun is not available to you. ## The decision, honestly framed Ask two questions. **How much irreducible night paging is there?** Count the pages that actually arrive outside business hours over a representative period, and separate them into noise, symptoms of fixable reliability problems, and genuine events that require a human at that hour. Most teams find the first two categories dominate. Fixing them is far cheaper than a second continent, and it is work that pays off during the day too. Only the third category is an argument for follow-the-sun. **Does the organization already span regions?** If there is already a real engineering team elsewhere that could plausibly own this service, follow-the-sun is a scheduling and training change — expensive but tractable. If there is not, this is a hiring and org-design decision with a multi-year horizon, and it should be argued on those terms, not smuggled in as an on-call improvement. ## When the answer is yes Follow-the-sun earns its cost when night paging is genuinely irreducible — a global user base with no quiet hours, or availability commitments that do not permit the recovery latency of a business-hours response — and when the second site can be a peer rather than a relay. It is also the honest answer when a single-region team has been quietly burning out for years and every cheaper lever has already been pulled. ## The cheaper alternatives worth naming first A partner team in another timezone taking a defined slice of hours, without full follow-the-sun ownership. Narrowing the set of conditions that page overnight to those that genuinely threaten the reliability target. Investing in the reliability work that makes the night quiet in the first place. Or a compensated, properly time-shifted night rotation within one region. Each of these is a fraction of the cost, and a strong answer walks the interviewer through them before reaching for a second site.

  • You have two sites and an incident is still open when the first site's day ends. How should that be handled?
    With an explicit live transfer rather than an implicit one: the incoming site joins while the outgoing responders are still present, receives the current hypothesis, mitigations in place and next action, and the ownership of the response is stated openly. A site should never discover at the start of its day that it inherited an incident nobody described.
  • How do you keep the second site from becoming a relay that escalates everything back?
    Give it real ownership, not just the pager — service ownership on paper, authority to change and deploy, access everywhere, and enough incidents to stay sharp. Measure the share of incidents that site resolves without escalating; if that number stays low, the model is not working and adding more process will not fix it.
  • What would you measure before proposing follow-the-sun to leadership?
    Out-of-hours page volume over a representative period, split into noise, fixable reliability issues and genuinely irreducible events; how many of those needed a human within minutes rather than hours; attrition and sick-leave signals on the current rotation; and the cost of the cheaper alternatives already tried. The proposal should stand on the irreducible remainder alone.

saying these in an interview costs you the question

  • Follow-the-sun just means adding a timezone to the schedule
  • Two engineers abroad are enough to cover nights
  • Handoffs are free if there is a written document
  • It removes the need to fix noisy alerts
  • The remote site can escalate back when it is stuck

context