Your only continuity plan for a live stream is last night's file backup of the broker data volumes. What has that already cost you by morning?
answer
- a file is a point, a stream moves
- the gap is the backup interval
- hours behind, not seconds behind
- several machines, several instants
- reader progress is not in the volumes
basics
~20 sA file backup fixes a stream at the instant it was taken, so every record written since is gone, along with every reader's progress. Streams keep moving while files do not, which is why the gap is counted in hours.
solid answer
~50 sA **file backup** is a copy of the broker data volumes taken from outside the cluster, on a schedule. It captures the stream at one instant, and the stream has been appended to ever since, so by morning the gap between the backup and reality is the whole night's records. Nothing fills that gap later: the records only ever existed on the cluster you lost, so a restore returns a stream that has jumped backwards in time. The bookkeeping around the records is a second loss — where each reader group had got to is not something the volume copy meaningfully preserves. The honest comparison is a second cluster kept fed by an ongoing cross-cluster copy, which trails the source by seconds of copy lag instead of hours, because it is fed by records as they are written rather than by a nightly job.
go deeper
Recall the one-line shape: a file backup captures a moment, a stream keeps growing, so the gap is however long since the last run. Say it in hours without hedging.
Explain the mechanism, not just the gap: the copy runs on a clock and outside the cluster's protocol, so it is both stale and stitched from several per-machine instants that need not agree.
Show judgment by separating the two jobs. Name the backup as the only artefact holding a known past state, and an ongoing cross-cluster copy as the only thing that keeps you serving, and refuse to let a plan substitute one for the other.
The argument to make is about what the organisation will actually pay for and rehearse. A nightly file copy is cheap and almost never used; a second fed cluster costs continuously and is the thing that answers a lost site. Say which risk you are buying down.
## What a file backup of a cluster actually is A **file backup** in this context means a copy of the cluster's data taken from *outside* the cluster: a filesystem copy or a volume snapshot of the disks the brokers write records to, together with whatever files hold the cluster's own bookkeeping. It is produced by a scheduler on a cadence, by a process that is not a participant in the cluster's protocol. Every property discussed below follows from those two facts — it runs on a clock, and it does not know what the cluster is doing while it runs. A stream is not a document. It is an append-only sequence that is still being appended to while the copy is taken, read by reader groups that are simultaneously advancing their **stored read positions**. Copying its files is like copying a ledger that someone is still writing in. ## The gap is measured in hours, not seconds The first and largest cost is pure staleness: 1. The backup fixes the stream's contents at the instant the copy began. 2. Writers keep appending for the rest of the night — that is the point of a live stream. 3. Nothing reconciles the two afterwards. The records written after the copy existed only on the cluster that was lost. So the answer to "what has it cost you by morning" is: **up to a full backup interval of records, plus a night of reader progress**. Shortening the interval shrinks that number but never removes it, because the interval is the whole mechanism. This is the deflating comparison that interviewers are fishing for. A second cluster kept fed by an ongoing **cross-cluster copier** trails the source by *seconds of copy lag*, because it is driven by records arriving rather than by a clock. A nightly file backup is three or four orders of magnitude worse on exactly the axis that matters, and teams routinely write it down as though the two were interchangeable. ## The second problem: the image is not internally consistent Even at the moment it is taken, a file backup of a multi-machine cluster is usually not one coherent picture: - The volumes are captured at slightly different instants, so a record can be present in one machine's files and absent from another's. - Writes are in flight during the copy — accepted by the cluster, not yet durably in every file the copy is reading. - The cluster's own bookkeeping (which stream lives where, which copies the cluster believed were caught up, what generation of leadership was in force) is captured at yet another instant, so a restored cluster can believe things about its data that the restored data does not support. What this looks like in practice varies by platform, and an interview answer that ignores that variation is describing one product: - Where each broker owns local files, you are stitching several per-machine instants into one cluster and hoping they agree. - Where the cluster writes into shared remote storage instead, the record files are not the fragile part — the cluster's metadata is, and a copy of it that disagrees with the storage is just as broken. - On a rented cluster there are usually no volumes to copy at all; the provider's own recovery mechanism is the only one you have. ## What the two approaches are actually for | | a file backup | a standby cluster fed by an ongoing copy | |---|---|---| | freshness | as old as the last scheduled run | seconds of copy lag behind the source | | what moves | files, at the storage layer | records, as they are written | | internal consistency | assembled from several per-machine instants | each record arrives as a record | | after a loss | rebuild, then restart everything against it | the other site is already running | | genuinely protects against | a past state you want back | losing a site | ## Where a file backup still earns its keep It is not worthless, and saying so is the mark of a considered answer rather than a memorised one. A file backup is the only artefact that holds a *known past state*. An ongoing cross-cluster copy is faithful by design: a stream emptied or removed by a mistaken administrative change is copied away just as faithfully as real traffic. The file is the one thing that predates the mistake. So the correct framing is that the two answer different questions. The backup answers "what did this look like before we broke it"; the ongoing copy answers "where do we keep serving from now that this site is gone". A continuity plan that lists only the first has confused an archive with a place to run.
- If every broker's volumes are captured at the same scheduled minute, why is the result still not one consistent cluster image?Because "the same minute" is not the same instant. Each volume is captured a little apart, writes are in flight while the copies run, and the cluster's bookkeeping is captured at a third moment. The pieces can disagree about which records exist and which machine held them, so the restored cluster's beliefs and its data do not line up.
- When is a file backup of a cluster genuinely worth keeping?When you want a known past state rather than a place to keep serving. An ongoing cross-cluster copy is faithful by design, so a stream emptied or removed by a mistaken administrative change is reproduced on the far side too. The file predates the mistake. That is an archive, not a continuity plan, and it should be written down as one.
A file backup of a live stream is a photograph of a river. It tells you exactly where the water was at one moment, and the river has been somewhere else ever since.
saying these in an interview costs you the question
- We back up nightly, so the stream can be rebuilt from it
- Restoring the volumes puts every reader group back where it was
- A file backup and a standby cluster protect against the same failure
- Running the backup more often turns it into a continuity plan
- Records written after the backup can still be recovered from the file