Where does the contract store beside a cluster keep its own state, and why does that decide whether a restore of the brokers is usable?
answer
- the service has state of its own
- ask who backs that state up
- state inside the cluster it serves
- ordering: cluster first, then the store
basics
~20 sA contract store keeps state of its own, and where it lives decides the restore: on the cluster it serves, in another team's database, or in a checked-in definition. Only the last comes back for free.
solid answer
~50 sThe store beside a cluster is a service with durable state of its own, and that state lives in one of three places. Some keep it in a stream on the very cluster they serve - so it returns only with that cluster's data, and only once the cluster is serving again, which makes the recovery circular and order-dependent. Some keep it in a separate database owned by another team, so the plan silently inherits that team's schedule and recovery point, which may be far worse than the one promised for the stream. Some derive it from a checked-in definition the deployment applies, which is cheapest to reproduce - but only if nothing is registered at runtime, because runtime entries exist in no artifact. Before declaring it restored, check that it answers at the address clients are configured with and that its contents are no older than the records about to be touched.
go deeper
Remember that the service standing beside the cluster stores data too, and that bringing the brokers back does not automatically bring it back.
Be able to name the placements - a stream on the cluster it serves, a separate database, or a checked-in definition - and say what each implies once the cluster is running again.
Demonstrate the ordering and the ownership problem: who backs that state up, on whose schedule, and in what sequence the two recoveries must run for a producer to work at the end of it.
The angle is accountability. Two recovery points owned by two teams are one plan only if somebody has written down which number governs and has rehearsed the sequence between them.
## The question behind the question A continuity plan that says "restore the cluster" is describing one service. Producers depend on at least two. The second one - the **contract store**, the companion service a producer must reach before it may publish - is not a stateless proxy: it holds durable content that took months to accumulate, and that content has its own storage, its own owner and its own backup story. An interviewer asking this is checking whether you have ever looked past the brokers, because a plan that restores the cluster and nothing else hands you a healthy cluster no producer can use. ## The three places that state lives | placement | who owns its backup | what happens when the cluster is restored | |---|---|---| | a stream on the very cluster it serves | you do, by accident - it rides along with the cluster's own data | it returns with that data and not before, so the service must be started after the cluster is serving; if the stream was not in what you restored, the store comes up empty | | a separate database owned by another team | that team, on their schedule | you inherit their recovery point and their sequence; two numbers nobody compared become one plan the moment you depend on them | | derived from a checked-in definition the deployment applies | source control, and the deployment that reads it | it is reproduced rather than restored, which is the cheapest path - provided nothing was registered at runtime | The first row is the one worth dwelling on, because it is genuinely circular. The store's content is kept on the cluster the store gates writes to. Recover the cluster and the content is there; fail to recover that part of the cluster and the store starts with nothing, and every producer is refused at the first publish - with the brokers healthy and the dashboards green. ## What a usable restore actually requires 1. **Sequence.** Whatever holds the state must be serving before the service that reads it starts. Where the state is on the cluster, that means the cluster first, then the store, then producers - and the plan has to say so, because nothing enforces the order. 2. **Reachability at the configured address.** A restored or standby instance answering on a different address is, to every client, an outage. This is the step most often skipped when the switch goes to a standby cluster rather than back to the original. 3. **Content no older than the records.** If the store's state was recovered to an earlier point than the stream's, producers and readers meet payloads whose contracts the store has never held. The store's own recovery point is therefore a number in your plan, not the database team's private business. 4. **A real publish and a real read.** Process-is-running is not restored. The proof is one write and one read completing through the restored path. ## Why the "rebuild it from source" answer is only half true Deriving the contents from a checked-in definition is the cheapest posture and the easiest to rehearse: no backup, no ordering, no second team. It is honest only where the definition is complete. Two things break it: - **runtime registration** - where producers add entries as they start or as they ship a new payload shape, those entries exist only in the store, and no artifact knows about them; - **drift** - the definition in source control is what somebody last committed, which is not necessarily what the running store holds. If you claim this posture, the honest sentence is "we reproduce it from source, and we have checked that nothing registers at runtime", not "we can rebuild it". ## What varies Not every estate has this dependency at all. Some platforms ship a contract store beside the cluster; others have none, and contracts travel with the payload or in a shared build-time artifact. Some designs also keep cluster metadata in a separate coordination service with its own state, while others keep it inside the cluster - and where that separate service exists, everything above applies to it too, with a wider blast radius. The way to phrase the answer so it survives any platform is to ask the general question: *for every service a client must reach besides a broker, where does its state live and who restores it?* ## The failure mode in one line A plan that covers the brokers and not the companion beside them does not half-work. It restores a cluster that refuses every write, which is indistinguishable, from the producer's side, from not having restored anything.
- In what order must the two recoveries run, and why?The cluster first, then the companion service, wherever the service keeps its state on that cluster - it has nothing to read until the cluster is serving. Where the state lives elsewhere the order is looser, but the service must be reachable and populated before the first producer is let back in.
- What do you check before declaring the companion service recovered?That it answers at the address clients are configured with, that its contents are no older than the records producers and readers are about to touch, and that one real publish and one real read complete through it. A running process is not evidence.
A spare key to the house, kept in a drawer inside that house. The copy genuinely exists, and it is genuinely useless until you are already inside - which is exactly the position a contract store is in when its own state lives on the cluster it serves.
saying these in an interview costs you the question
- Assumes restoring the brokers restores everything a producer needs.
- Believes the companion service holds no durable state of its own.
- Never asks who owns the backup of that state, or on what schedule.
- Thinks producers will simply re-register whatever the store lost.
- Treats the two restores as independent, with no ordering between them.