After moving to a hosted broker, a team proposes halving its broker on-call rota - which incident classes still page it?
answer
- rented work and paging work are different sets
- residual duties are the incident classes
- skills narrow, coverage does not
- who fixes a node, not who is paged
- implicit de-staffing is the real risk
basics
~20 sEvery incident whose cause is a residual duty: a stream that was never created or was misnamed, records aged out too early, a permission change that cut off a client, a client version mismatch, readers falling behind, and a bill nobody watched. Renting changed who repairs a node, not who is paged when delivery stops.
solid answer
~50 sThe rota does not shrink by the fraction of the work that was rented, because the two sets are different. What the provider took over - dead nodes, engine patches, the endpoint, hardware - were rarely the incidents that woke people; they are also the ones a provider is measurably good at. What survives is exactly the `residual duties`: a stream that does not exist or carries a name another team depended on, a retention choice that removed records before a consumer needed them, a permission or credential change that cut off a producer, a client version that stopped working, a reading side that has stopped keeping up, and spend that grew unnoticed. You can shrink the rota's required *skills* - nobody needs to be able to rebuild a node at three in the morning - but not its coverage. Renting changed who fixes a node; it did not change who is paged when delivery stops.
go deeper
The point to carry away is that a hosted broker does not watch your consumers, your permissions or your streams. Someone on your side still has to answer when messages stop arriving.
Explain why the two sets are disjoint: the rented duties were mostly scheduled infrastructure work, while the paging incidents come from stream design, retention, permissions, clients and the reading side - none of which transferred.
Give the concrete list of surviving incident classes and say precisely what you would cut instead: the deep repair skills and their drills, keeping full coverage. Interviewers are listening for whether you have actually been paged for a rented cluster.
The trade-off to own is implicit de-staffing. Each residual duty needs a named owner before the rota changes, and the classes that are now escalation rather than repair still need somebody awake with the evidence to escalate.
This is the question that separates people who have run a rented cluster from people who have read a feature page. The proposal sounds reasonable: the provider now does the operating, so the rota should shrink in proportion. It does not, and the reason is structural. ## Two different sets The duties a rental transfers and the incidents that actually page a team are largely **disjoint sets**. The transferred duties - replacing a dead node, patching the engine, keeping the endpoint answering, replacing failed disks - were real work, but they were scheduled work far more often than they were 3 a.m. work, and they are precisely what a provider operating thousands of clusters is good at. The incidents that wake people are overwhelmingly caused by the **residual duties**, and not one of those moved. ## The incident classes that survive renting 1. **The stream that does not exist, or has the wrong name.** A service deploys and publishes into nothing, or a consumer subscribes to a name that was quietly changed. Stream design and naming never transferred. 2. **Records that are gone before someone needed them.** The retention choice is yours; a replay or a late consumer discovers it. The provider enforced exactly the policy it was given. 3. **A permission or credential change that cut off a client.** Authorisation rules and credential lifetimes are residual. A rotation nobody coordinated produces a total outage for one producer and a perfectly healthy cluster. 4. **A client version that stopped working.** Your producing and reading services were not upgraded by the rental, and a provider's engine upgrade can change what old clients negotiate. 5. **The reading side stopping.** Where readers own a stored position, unread records pile up behind a stalled consumer group; on destructive-read designs, unacknowledged work stops draining and eventually redelivers. Either way the provider is watching its cluster, not your consumers - and this is the single most common surviving page. 6. **Spend that grew.** A design change multiplies what is stored or what flows, and the first notice is a bill. Nobody outside your team is looking. | Incident cause | Who repairs it after renting | |---|---| | Node hardware failure | Provider | | Engine defect fixed by a patch | Provider | | Endpoint unreachable | Provider | | Missing or misnamed stream | You | | Records aged out before use | You | | Permission or credential change | You | | Incompatible client version | You | | Reading side not keeping up | You | | Unexpected spend | You | ## What legitimately changes about the rota The honest reductions are real but narrower than the proposal assumes: - **The required skill set narrows.** Nobody needs to be able to rebuild a node, restore a volume or balance data across machines under pressure, because that capability is gone whether you want it or not. - **Some classes become escalation rather than repair.** When the cause is on the provider's side, your job is detection, evidence and escalation - not a fix. That is faster to train for, but it still needs somebody awake. - **The frequency of infrastructure pages drops.** Providers genuinely are better at fleet repair than a small team. What does not change is **coverage**. Someone still has to be reachable when delivery stops, because the residual classes above produce exactly the same symptom as an infrastructure failure - messages are not arriving - and you cannot tell which it is without looking. ## The organisational trap The failure mode is not the rota decision itself; it is **implicit de-staffing**. Each residual duty quietly loses its owner because "we don't run the broker any more", and then an incident lands on whoever is nearest. Before shrinking anything, write the residual duties down and put a name against each one. A duty nobody owns is not a duty that went away - it is a duty that will be discovered during an incident, by someone who did not know it was theirs. The defensible position in an interview: *renting changed who fixes a node, but not who is paged when delivery stops.* Say what you would actually cut - the deep infrastructure skills and the drills for them - and what you would keep at full coverage, and you have answered the question the way a senior engineer is expected to.
- Which on-call capability can you genuinely retire after moving to a hosted broker?The deep repair skills and the drills that kept them sharp: rebuilding a node, restoring a volume, rebalancing data across machines under pressure. That capability is gone whether you want it or not, so training for it is waste. Detection, evidence-gathering and escalation replace it, and they are cheaper to train.
- Why does a reading side falling behind remain the most common surviving page?Because it sits entirely in the residual set and produces the same symptom as an infrastructure failure - data not arriving downstream - while the provider's cluster stays healthy by every measure it publishes. Nobody outside your team is looking at your consumers, and no rental has ever included them.
- How would you argue against the proposal without arguing against the move to a hosted tier?Separate the two claims. Renting is a good trade on infrastructure work, and the savings are real in skills and in the frequency of hardware pages. The rota question is different: it is set by the residual duties, which did not move. Cut the drills, keep the coverage, and name an owner for each residual duty.
saying these in an interview costs you the question
- Assumes the provider's page replaces yours when delivery stops
- Scales the rota by the fraction of work that was rented
- Says nothing can page you if the provider's cluster is healthy
- Forgets that readers falling behind was never part of any rental
- Treats an unowned residual duty as a duty that disappeared
- Cuts coverage instead of cutting the deep repair drills