skip to content

Backup & Restore Reality

Why a live stream is seldom rebuilt from a file backup, what the copy in remote storage is genuinely for, and the metadata a restore forgets. Asked because positions and permissions come back empty.

part ofBroker & streaming operationsoverview, primer and where to startread it →
on this pageshow

questions

3

Your only continuity plan for a live stream is last night's file backup of the broker data volumes. What has that already cost you by morning?

level: juniorimportance: must knowfreq 58%

answer

  1. a file is a point, a stream moves
  2. the gap is the backup interval
  3. hours behind, not seconds behind
  4. several machines, several instants
  5. reader progress is not in the volumes

basics

~20 s

A file backup fixes a stream at the instant it was taken, so every record written since is gone, along with every reader's progress. Streams keep moving while files do not, which is why the gap is counted in hours.

solid answer

~50 s

A **file backup** is a copy of the broker data volumes taken from outside the cluster, on a schedule. It captures the stream at one instant, and the stream has been appended to ever since, so by morning the gap between the backup and reality is the whole night's records. Nothing fills that gap later: the records only ever existed on the cluster you lost, so a restore returns a stream that has jumped backwards in time. The bookkeeping around the records is a second loss — where each reader group had got to is not something the volume copy meaningfully preserves. The honest comparison is a second cluster kept fed by an ongoing cross-cluster copy, which trails the source by seconds of copy lag instead of hours, because it is fed by records as they are written rather than by a nightly job.

go deeper

for a junior

Recall the one-line shape: a file backup captures a moment, a stream keeps growing, so the gap is however long since the last run. Say it in hours without hedging.

for a middle

Explain the mechanism, not just the gap: the copy runs on a clock and outside the cluster's protocol, so it is both stale and stitched from several per-machine instants that need not agree.

for a senior

Show judgment by separating the two jobs. Name the backup as the only artefact holding a known past state, and an ongoing cross-cluster copy as the only thing that keeps you serving, and refuse to let a plan substitute one for the other.

for a principal

The argument to make is about what the organisation will actually pay for and rehearse. A nightly file copy is cheap and almost never used; a second fed cluster costs continuously and is the thing that answers a lost site. Say which risk you are buying down.

## What a file backup of a cluster actually is A **file backup** in this context means a copy of the cluster's data taken from *outside* the cluster: a filesystem copy or a volume snapshot of the disks the brokers write records to, together with whatever files hold the cluster's own bookkeeping. It is produced by a scheduler on a cadence, by a process that is not a participant in the cluster's protocol. Every property discussed below follows from those two facts — it runs on a clock, and it does not know what the cluster is doing while it runs. A stream is not a document. It is an append-only sequence that is still being appended to while the copy is taken, read by reader groups that are simultaneously advancing their **stored read positions**. Copying its files is like copying a ledger that someone is still writing in. ## The gap is measured in hours, not seconds The first and largest cost is pure staleness: 1. The backup fixes the stream's contents at the instant the copy began. 2. Writers keep appending for the rest of the night — that is the point of a live stream. 3. Nothing reconciles the two afterwards. The records written after the copy existed only on the cluster that was lost. So the answer to "what has it cost you by morning" is: **up to a full backup interval of records, plus a night of reader progress**. Shortening the interval shrinks that number but never removes it, because the interval is the whole mechanism. This is the deflating comparison that interviewers are fishing for. A second cluster kept fed by an ongoing **cross-cluster copier** trails the source by *seconds of copy lag*, because it is driven by records arriving rather than by a clock. A nightly file backup is three or four orders of magnitude worse on exactly the axis that matters, and teams routinely write it down as though the two were interchangeable. ## The second problem: the image is not internally consistent Even at the moment it is taken, a file backup of a multi-machine cluster is usually not one coherent picture: - The volumes are captured at slightly different instants, so a record can be present in one machine's files and absent from another's. - Writes are in flight during the copy — accepted by the cluster, not yet durably in every file the copy is reading. - The cluster's own bookkeeping (which stream lives where, which copies the cluster believed were caught up, what generation of leadership was in force) is captured at yet another instant, so a restored cluster can believe things about its data that the restored data does not support. What this looks like in practice varies by platform, and an interview answer that ignores that variation is describing one product: - Where each broker owns local files, you are stitching several per-machine instants into one cluster and hoping they agree. - Where the cluster writes into shared remote storage instead, the record files are not the fragile part — the cluster's metadata is, and a copy of it that disagrees with the storage is just as broken. - On a rented cluster there are usually no volumes to copy at all; the provider's own recovery mechanism is the only one you have. ## What the two approaches are actually for | | a file backup | a standby cluster fed by an ongoing copy | |---|---|---| | freshness | as old as the last scheduled run | seconds of copy lag behind the source | | what moves | files, at the storage layer | records, as they are written | | internal consistency | assembled from several per-machine instants | each record arrives as a record | | after a loss | rebuild, then restart everything against it | the other site is already running | | genuinely protects against | a past state you want back | losing a site | ## Where a file backup still earns its keep It is not worthless, and saying so is the mark of a considered answer rather than a memorised one. A file backup is the only artefact that holds a *known past state*. An ongoing cross-cluster copy is faithful by design: a stream emptied or removed by a mistaken administrative change is copied away just as faithfully as real traffic. The file is the one thing that predates the mistake. So the correct framing is that the two answer different questions. The backup answers "what did this look like before we broke it"; the ongoing copy answers "where do we keep serving from now that this site is gone". A continuity plan that lists only the first has confused an archive with a place to run.

  • If every broker's volumes are captured at the same scheduled minute, why is the result still not one consistent cluster image?
    Because "the same minute" is not the same instant. Each volume is captured a little apart, writes are in flight while the copies run, and the cluster's bookkeeping is captured at a third moment. The pieces can disagree about which records exist and which machine held them, so the restored cluster's beliefs and its data do not line up.
  • When is a file backup of a cluster genuinely worth keeping?
    When you want a known past state rather than a place to keep serving. An ongoing cross-cluster copy is faithful by design, so a stream emptied or removed by a mistaken administrative change is reproduced on the far side too. The file predates the mistake. That is an archive, not a continuity plan, and it should be written down as one.

A file backup of a live stream is a photograph of a river. It tells you exactly where the water was at one moment, and the river has been somewhere else ever since.

saying these in an interview costs you the question

  • We back up nightly, so the stream can be rebuilt from it
  • Restoring the volumes puts every reader group back where it was
  • A file backup and a standby cluster protect against the same failure
  • Running the backup more often turns it into a continuity plan
  • Records written after the backup can still be recovered from the file
open as a page

A broker cluster is restored from a file backup and the records are present, yet no reader group progresses and clients are refused. What did the restore not return?

level: seniorimportance: must knowfreq 53%

basics

~20 s

The bookkeeping around the records: stored read positions, the entries saying who may connect and act on which stream, per-stream settings that differed from the cluster defaults, and the contract store the payload identifiers resolve against. The records are the easy part.

open as a page

Your cluster already parks closed history in cheap object storage, and the continuity plan calls that the backup. What is that remote-storage copy genuinely for?

level: middleimportance: should knowfreq 41%

basics

~20 s

It exists to make long history affordable and readable through the cluster, not to be restored from. The parked objects are addressed through the cluster's own bookkeeping, so on their own they are opaque files with no reader positions, permissions or settings attached.

open as a page