What has to transfer at an on-call shift handoff, and what goes wrong when the handoff is just a message saying "quiet week"?
answer
- state, not a status summary
- everything muted needs an expiry
- in-flight changes outlive the shift
- temporary fixes need a ticket and an owner
- a written record plus short live overlap
basics
~20 sA handoff transfers state, not a summary: open incidents and their next step, every alert silenced during the shift with its expiry, in-flight changes and migrations, temporary manual fixes with an owner, and risky events scheduled in the next shift.
solid answer
~50 sI treat the handoff as a transfer of live state rather than a status report. The outgoing person hands over: any open or recently mitigated incident with the current hypothesis and the next action; every silence or suppression created during the shift, each with an expiry and a reason; in-flight changes — a paused rollout, a running migration, a flag left at partial exposure; any manual mitigation still holding the system up, such as a cluster scaled by hand or a job disabled; and anything risky scheduled into the incoming shift. The failure mode of "quiet week" is that muted alerts and manual fixes survive the shift invisibly, so the next occurrence never pages and the temporary workaround becomes permanent unowned state. Practically I want a written handoff in the team's channel or ticket plus ten minutes of live overlap, because the written form is the record and the call is where the incoming person asks the questions.
go deeper
Know that a handoff is a checklist you run at the end of every shift, quiet or not, and that anything you silenced or changed by hand during the night must be written down before you sign off.
Be able to list the items that transfer and explain why each one bites: open silences hide recurrences, in-flight changes produce confusing symptoms, and manual mitigations drift away from the declared configuration.
Demonstrate the enforcement side — silences that expire by construction, tickets opened for every temporary fix before the shift ends, and a live transfer when an incident is still open rather than a written note.
Own the format across teams: a handoff artefact that a postmortem can read months later, expectations that hold when shifts cross sites or timezones, and a way to detect teams whose handoffs have quietly decayed into "all quiet".
## The handoff transfers state, not a summary At the end of a shift the outgoing engineer holds a set of facts that exist nowhere else: what they silenced, what they changed by hand, what they were about to investigate. A handoff is the moment that state either moves to the next person or disappears. "Quiet week, good luck" moves nothing. ## The checklist worth memorising **Open and recently closed incidents.** Not just that there was one — the current hypothesis, what was already tried and ruled out, what mitigation is in place, whether it is a real fix or a hold, and the single next action. An incident closed two hours ago is still live information, because the same failure often returns within a day. **Every silence, suppression or muted alert, with an expiry.** This is the item that most often causes a repeat incident. Someone mutes a flapping alert at 2am to get an hour of sleep, the mute has no end time, and three weeks later the real occurrence of that condition pages nobody. The rule worth introducing is that no silence outlives the shift that created it: every one gets an expiry no later than the handoff and a reason, and anything that needs to stay muted longer needs a ticket and an explicit owner. **In-flight changes.** A rollout paused halfway, a schema migration running in the background, a backfill job, a feature flag sitting at partial exposure, a frozen pipeline. Each of these can page the incoming person with a symptom that is meaningless without knowing a change is mid-flight. **Manual mitigations still holding the system up.** A node pool scaled up by hand, a cron disabled, rate limits loosened, a replica promoted. These are invisible to anyone who did not do them, they drift away from the declared configuration, and they get reverted by the next deploy at the worst moment. Every one needs a line in the handoff plus a ticket to either unwind it or make it real. **Degraded dependencies and open vendor tickets.** If an upstream provider is in a partial outage, the incoming person needs the ticket number and the expected update time, otherwise they debug a system that is not broken. **What is scheduled into the next shift.** A launch, a database maintenance window, a load test, a planned failover. Whoever is holding the pager should not learn about these from the alert. ## Written record plus a short live overlap The written handoff is the artefact — it is searchable, it survives, and it is what a postmortem reads later. The live overlap of ten or fifteen minutes is where the incoming person asks the question that was not written down and confirms they can actually be reached. Teams that do only the call lose the record; teams that do only the document find that nobody reads it during a busy week. Doing both is cheap. ## Handing over an incident that is still open Avoid it where you can — a shift boundary in the middle of an active incident is a documented moment of risk. When it is unavoidable, do it live rather than in writing: the incoming person joins the call while the outgoing person is still on it, hears the current state, and the transfer of who is responsible is stated explicitly and out loud so everyone on the bridge knows who to answer to. The role-transfer mechanics of incident command are their own subject; what belongs to the rotation is the principle that the pager and the responsibility move together, and never silently. ## Why the discipline pays Almost every one of these items is a mechanism by which one shift's convenience becomes the next shift's incident. Silences hide recurrences, manual fixes decay, in-flight changes produce confusing symptoms. A handoff format that names each of them turns a personal memory into shared state, and it costs about ten minutes a shift.
- Someone silenced a flapping alert at 2am and left it open-ended. What rule would you introduce?No silence outlives the shift that created it. Every silence gets an expiry no later than the handoff, a reason, and a named owner; anything that must stay muted longer needs a ticket saying why and when it will be fixed. Otherwise the muted condition's next real occurrence pages nobody, and the alert quietly stops existing.
- A high-severity incident is still open when your shift ends. How do you hand it over?Live, not in writing. The incoming person joins the call while I am still on it, I state the current hypothesis, what has been ruled out, what mitigation is holding, and the next action, and the transfer of responsibility is said out loud so everyone knows who owns the response now. I stay reachable for a while afterwards.
- What belongs in the written handoff versus the live overlap?Written: the durable state — silences with expiries, in-flight changes, manual mitigations and their tickets, degraded dependencies, upcoming risky events. Live: the judgment that is hard to write, such as which symptom to distrust and what the outgoing person suspects but could not prove, plus confirming the incoming person can actually receive pages.
saying these in an interview costs you the question
- A quiet shift does not need a handoff
- Silences can stay until someone notices
- The incoming person can read the alert history themselves
- Handoff just means the pager is yours now
- A manual fix does not need a ticket if it works