A broker cluster is restored from a file backup and the records are present, yet no reader group progresses and clients are refused. What did the restore not return?
answer
- records are back, bookkeeping is not
- positions, permissions, settings, contracts
- an empty position falls to a default
- test with an unprivileged client
- declared state reapplies, hand-made state does not
basics
~20 sThe bookkeeping around the records: stored read positions, the entries saying who may connect and act on which stream, per-stream settings that differed from the cluster defaults, and the contract store the payload identifiers resolve against. The records are the easy part.
solid answer
~50 sA restore brings back **records**; what makes a cluster useful is the state wrapped around them, and that is where a file restore comes back empty. Four things go missing in practice. **Stored read positions**: on platforms where reader progress lives in the cluster, it comes back either stale or absent, so a reader group falls to a default start point and either reprocesses everything or silently skips the gap — and on destructive-read platforms there is no stored position at all, only an undelivered set whose delivery state is equally unreliable. **Permission entries**: often held outside the volumes you copied, which is exactly why clients are being refused. **Per-stream settings** that differed from the cluster default: the stream comes back inheriting defaults, quietly. **The contract store**: payload identifiers now resolve against nothing. Rehearse a restore and you find these in an afternoon; skip it and you find them during the incident.
go deeper
Hold on to the headline: a restore returns records, not the state around them. Reader progress, permissions, settings and payload contracts are separate things that may not come back.
Be able to list the four omissions and give the consequence of each — where a reader group starts when it finds no stored position, why clients are refused, why a stream quietly keeps less history, why payloads cannot be interpreted.
Show that you would prove it rather than assume it: restore into a scratch cluster, connect with an unprivileged client, start a positionless reader group, compare settings, interpret one real payload. Name which parts differ between platforms.
The leverage is upstream. Argue for surrounding state that is declared and reproducible rather than accumulated by hand, so a restore is a reapply instead of an archaeology exercise, and budget the rehearsal that turns this list from theory into a checklist.
## The records are the easy part A restore from a file backup is usually judged on the wrong question — *are the records there?* — and that question almost always answers yes. The operational failure comes from the state that surrounds the records, which is held in different places, owned by different systems, and rarely captured by whatever copied the broker data volumes. The symptom in the question is the standard shape: data present, nobody working. Walk the four omissions in order and it resolves quickly. ## The four things that come back empty 1. **Reader progress.** On platforms where a reader group's **stored read position** lives in the cluster, a restore returns it as of the backup instant at best, and absent at worst. A reader group that finds no stored position falls back to a default start point. Which default varies — oldest or newest — and both are bad in a different way: from the oldest, every record is reprocessed, which is a duplicate-work event sized like your whole retention; from the newest, everything between the backup and now is silently skipped and nobody gets an error. On a destructive-read platform there is no stored position to restore at all; what matters there is the set of undelivered records and their delivery state, which a volume copy is no better at preserving. 2. **Who may do what.** The entries that say which principal may connect and act on which stream frequently live outside the data volumes — in a separate store, an external directory, or the platform's control plane. Restore the data and those entries can be empty. Empty fails one of two ways depending on the platform's posture: everything is refused (your symptom) or, far worse, nothing is. 3. **Per-stream settings.** Streams that had been given something other than the cluster default — a different history length, a different copy count, a different cleanup behaviour — come back inheriting the default if the settings were not in what you copied. This one is quiet. The cluster is up, the readers are reading, and a stream is now keeping far less history than the team believes. 4. **The contracts the payloads reference.** Records generally carry an identifier that is resolved against a **contract store** standing beside the cluster. Restore the records without it and you have bytes that no consumer can interpret. This is the omission that makes people realise a broker restore is a *system* restore. ## Reading the symptom | symptom after the restore | the omission it points at | |---|---| | clients cannot connect or are refused an action | the permission entries were not in the copy | | readers start from the very beginning | no stored read position, default is the oldest | | readers are live but the gap is never processed | no stored read position, default is the newest | | history disappears sooner than expected | the stream fell back to the cluster default | | readers connect and then fail on every payload | the contract store did not come with the cluster | ## Why this is discovered during the incident The backup job reports success because it copied what it was pointed at. Nothing in that report knows about a permission store on another machine, a control plane that owns the settings, or a companion service holding the contracts. The gap is invisible to the only signal anybody watches. The cure is unglamorous and cheap: - Restore into a scratch cluster on a normal working day — not as a drill you schedule annually and then cancel. - Connect with an **ordinary, unprivileged client**, not an administrative one, because an administrative client hides the entire second omission. - Start a reader group that has no stored position and watch where it begins. - Compare the restored streams' settings against what the teams believe they have. - Try to interpret one real payload end to end. Each step maps to one row of the table above, and the whole exercise fits in an afternoon. ## What varies between platforms Say this out loud in an interview, because a universal claim here is wrong on some real system: - Where reader progress is stored differs — inside the cluster, in the client's own storage, or nowhere at all because reads are destructive. - Where permission entries live differs — with the cluster's data, in a separate service, or in a provider's identity system entirely outside your backup. - Where per-stream settings live differs — with the data, in a metadata service, or in the infrastructure definition that created the stream, which is the one case where a restore is genuinely easy because the definitions can simply be reapplied. That last case is the constructive end of this answer. The omissions hurt least when the surrounding state is declared somewhere reproducible rather than accumulated by hand on a running cluster: reapplying declarations after a restore is a minute of work, whereas reconstructing what nobody wrote down is the incident.
- Why is a reader group that finds no stored read position after a restore dangerous in two opposite ways?Because it falls to a default start point, and which default the platform uses decides the damage. Starting from the oldest record reprocesses the entire retained history, a duplicate-work event sized like your retention. Starting from the newest skips everything between the backup and now, silently, with no error anywhere. You cannot reason about the restore until you know which one your platform does.
- The restore finished and an administrator can read every stream. Why does that prove almost nothing?Because an administrative client typically bypasses the permission entries that ordinary clients depend on, so the single largest omission is invisible to it. The test has to be an ordinary, unprivileged client acting as a real application would. Until one of those connects, reads and writes successfully, the restore has only been shown to work for the one identity that never needed the entries.
- Which of these omissions is quiet rather than loud, and why does that make it worse?Per-stream settings. Missing permissions refuse connections and missing contracts fail every payload — both are loud and get fixed within minutes. A stream that silently fell back to the cluster default keeps working perfectly while holding far less history than its owners believe, and that is discovered only when someone tries to read back further than it now goes.
saying these in an interview costs you the question
- The records came back, so the restore succeeded
- Reader groups resume exactly where they stopped after a restore
- Permission entries are always inside the broker data volumes
- A restored stream keeps the settings it had before, automatically
- An administrator reading the restored streams proves clients will work
- Payload contracts are embedded in the records, so nothing else is needed