skip to content

Which services standing beside your clusters belong in the continuity plan, and what do you do about the ones you exclude?

level: principalimportance: should knowfreq 32%

answer

  1. list what a publish actually touches
  2. each companion gets a posture
  3. shared instance, per-site copy, or rebuild
  4. the weakest part governs the plan

basics

~20 s

Inventory everything a publish and a read touch besides the brokers, then give each companion a posture: one shared instance, a copy per site, or rebuild on demand. Anything excluded should be excluded on purpose, in writing.

solid answer

~50 s

Start with an inventory: for one publish and one read, list everything the client touches that is not a broker. Then give each entry a posture. A single shared instance is cheap and keeps contents consistent, but it makes the site hosting it a single point of failure for every other site. A copy per site removes that and buys a second copying problem - two instances that drift, with nobody noticing until a switch. Rebuild on demand is cheapest of all, and honest only where the contents come from a checked-in definition and the rebuild fits inside the recovery time you committed to. Two rules survive every estate: the plan's effective recovery point is the worst of its parts, not the cluster's; and the evidence that counts is a rehearsal in which clients were pointed at the standby and a publish succeeded.

go deeper

for a junior

Know that a continuity plan covers more than the brokers, and that the services standing beside the cluster have to appear in it explicitly rather than by assumption.

for a middle

Explain the three postures available for each companion - one shared instance, a copy per site, or rebuild on demand - and what each of them costs to run.

for a senior

Show how you would test the plan: point clients at the standby, publish, read, and see whether anything outside the cluster had to be fixed by hand before either worked.

for a principal

Own the trade: correlated failure from a shared instance against cost and drift from one per site, and the effective recovery point being the worst of the parts rather than the cluster's.

## Start with an inventory, not with a policy The recurring failure is not choosing the wrong posture for a companion service - it is never having listed it. So the first move is mechanical: take one publish and one read, and write down everything the client touches that is not a broker. On most estates the list is short and uncomfortable: - the **contract store**, which a producer must reach before it may publish; - on designs that keep cluster metadata outside the brokers, the **coordination service** that holds it; - the service that answers whether a client may connect at all; - whatever resolves the address the client connects to. Each entry on that list is a service that can be healthy while the cluster is gone, gone while the cluster is healthy, or absent from a recovered site entirely. None of them is optional at the moment of a switch. ## Three postures, and what each one really costs | posture | cost | how it fails | when it is the right call | |---|---|---|---| | one shared instance, hosted at one site | lowest; one thing to run and monitor | losing that site stops publishing at the sites that survived, which is the opposite of what the estate was built for | small estates, or a companion whose loss is genuinely tolerable for the length of the outage | | a copy per site | highest; two deployments, and a second copying problem between them | the two instances drift, and nobody notices until a switch needs an entry the local one never received | the companion gates writes and the estate genuinely intends to keep serving from either site | | rebuild on demand from a checked-in definition | low, and rehearsable in a pipeline rather than in an incident | silently incomplete wherever anything is registered at runtime | contents are fully derived from source and the rebuild fits inside the recovery time | The middle row is the one that gets chosen on a slide and abandoned in practice. Keeping two instances of a companion in step is its own copying problem, with its own lag and its own failure modes, and it is invisible until the day it matters. ## The number that actually governs A continuity plan has one honest recovery point, and it is the worst of the parts a producer needs - not the one written next to the cluster. If the stream is copied continuously and its contract store is recovered from a nightly export, the plan's recovery point is the nightly one for every payload whose contract was added since. The same arithmetic applies to recovery time: the wait is the cluster's start, plus the companion's, plus whatever hand-editing stands between them. This is the claim worth making out loud in an interview, because it reframes the whole conversation from "is the cluster covered" to "is the publish path covered". ## What a lead actually signs off 1. **The list**, with an owner per entry - not the messaging team by default, because several of these belong to somebody else. 2. **A posture per entry**, with its cost accepted rather than assumed. 3. **The sequence**, because most of these have an order: something has to be serving before the next thing can read it. 4. **An exclusion register.** A dependency left out deliberately, with the consequence written down, is a decision. One left out by omission is the outage. 5. **A rehearsal date.** Pointing clients at the standby and completing a publish is the only test that exercises the companions at all; a backup job reporting success does not. ## What varies, and why the answer must say so Not every estate has the same companions. Some platforms ship a contract store beside the cluster and some have none, in which case contracts ride with the payload or come from a shared build-time artifact and there is nothing to plan for. Some designs keep cluster metadata in a separate coordination service and others inside the cluster itself. A rented cluster hides several of these behind the provider's boundary, which changes the question from "how do I recover it" to "what has the provider promised about it, and is that number worse than mine?" - and a promise you cannot rehearse is a posture too, just not a comfortable one. ## The honest closing position Most organisations will pay for one shared companion instance and a rehearsal once a year, and that is a defensible answer if it is said out loud with its consequence attached: losing the hosting site stops publishing everywhere until it is rebuilt. The indefensible version is the same arrangement described as full cross-site continuity, because the brokers were the only thing anyone counted.

  • What makes a copy of a companion service per site harder than it sounds?
    Two instances drift. Anything added at one site has to reach the other, which is a second copying problem with its own lag and failure modes, and nobody notices the divergence until a switch - when producers at the surviving site need entries the local instance never received.
  • Which companion can honestly be left out of the plan?
    One whose contents are fully reproducible from a checked-in definition the deployment applies, and whose rebuild fits inside the recovery time you committed to. Write the reasoning down: a dependency excluded on purpose is a decision, and one excluded by omission is the outage.

saying these in an interview costs you the question

  • Puts only the brokers in the continuity plan.
  • Assumes a shared companion at one site cannot cause a correlated failure.
  • Counts a dependency as covered because it is named in a runbook.
  • Leaves a companion with a weaker recovery point than the stream it gates.
  • Believes a copy per site removes the work of keeping the two in step.