skip to content

Your cluster already parks closed history in cheap object storage, and the continuity plan calls that the backup. What is that remote-storage copy genuinely for?

level: middleimportance: should knowfreq 41%

answer

  1. it belongs to the cluster, not beside it
  2. cheap length, not a recovery artefact
  3. the index lives in cluster metadata
  4. only closed history is parked
  5. removing the stream takes it too

basics

~20 s

It exists to make long history affordable and readable through the cluster, not to be restored from. The parked objects are addressed through the cluster's own bookkeeping, so on their own they are opaque files with no reader positions, permissions or settings attached.

solid answer

~50 s

The **remote-storage copy** is closed history that the cluster itself has parked in cheap object storage so that keeping a long stream does not mean buying a long row of local disks. It is part of the live cluster's storage, and that is the point people miss: readers reach it *through* the cluster, which knows which object holds which part of which stream. Lose the cluster and its bookkeeping and you have a bucket of files whose layout is the platform's private business, with nothing in them that says where a reader group had got to, who was allowed to read it, or what each stream was configured to do. Its lifecycle is also the stream's lifecycle on most platforms — remove the stream and its parked history goes with it, so it is not a second point in time either. Write it in the plan as cheap long history, not as a restore route.

go deeper

for a junior

Remember the one distinction: parked history is about keeping a long stream cheaply, not about getting a dead cluster back. Durable storage underneath does not make it a recovery plan.

for a middle

Explain the dependency. The objects are located and interpreted through the cluster's own metadata, only closed history is out there, and the newest records — the ones an incident is about — are still on the machines.

for a senior

Demonstrate that you would test the claim instead of debating it: stand up a fresh cluster against that storage and see whether a single record can be served. Say plainly that the answer differs between platforms.

for a principal

The call to own is where each line of the plan belongs. Cheap long history is a storage decision with a cost curve; surviving a lost site is a separate purchase, and conflating them is how an organisation discovers it bought neither.

## What the remote-storage copy is The **remote-storage copy** is closed history that the cluster has written out to cheap object storage rather than keeping on the machines' own disks. It is a storage-tiering arrangement made by the cluster, for the cluster. Two consequences follow immediately, and both are what an interviewer is checking: - it is *inside* the system, not beside it — the cluster is the thing that knows what those objects are; - it was created to change the **cost of holding history**, not to create an independent recovery artefact. That second point is the whole question. Because the storage is durable and off-machine, it *looks* like a backup, and continuity plans keep recording it as one. ## What it genuinely buys you - **Affordable length.** History that would otherwise be capped by local disk can be kept far longer for a fraction of the price. - **Reach for readers.** A reader that has fallen a long way behind, or one that is deliberately reading old history, can still be served, because the cluster fetches the parked objects on its behalf. - **Smaller, cheaper machines.** Storage stops being the thing that sizes the cluster, so capacity decisions are driven by throughput instead. - **Less pressure on the local volumes**, which is an operational win in its own right. None of those is a recovery property. Every one of them is a property *of the running cluster*. ## Why it is not a restore route 1. **The objects are addressed through the cluster's bookkeeping.** Which object holds which span of which stream is metadata the cluster owns. That mapping is the platform's private business, it differs between platforms, and it is not in the objects in a form you are meant to reconstruct by hand. 2. **It holds only closed history.** The newest records — precisely the ones an incident is about — are by construction still on the cluster and not yet parked. 3. **Its lifecycle is the stream's lifecycle.** On most platforms, removing a stream removes its parked history too, and the deletion rules that expire records apply through the cluster. A mistaken removal is therefore not survived by the parked objects the way people assume. 4. **It carries no bookkeeping at all.** No stored read positions, no permission entries, no per-stream settings, nothing about the contracts the payloads reference. Even a perfect object-for-object recovery gives you records and nothing around them. ## Two copies, two jobs | | the remote-storage copy | a file backup taken from outside | |---|---|---| | who wrote it | the cluster, as part of its storage | a scheduler, outside the cluster | | what it holds | closed history only | whatever was on the volumes at one instant | | read by | readers, through the cluster | nobody, until it is restored | | independent of the cluster | no — addressed by cluster metadata | yes, but stale and stitched from instants | | survives removing the stream | usually not | yes, that is its one real strength | ## What varies, and how to say so Be explicit that platforms differ here, because a confident universal claim is the defect this subject invites: - Some platforms keep enough index alongside the parked objects that a rebuilt cluster can adopt them; others keep that index only in cluster metadata. - Some let a client read parked history directly for bulk processing; others expose it exclusively through the broker path. - On a rented cluster the whole arrangement may be invisible — you are buying a retention length, and there is no bucket of yours to inspect. The safe formulation is the one that survives all three: *the remote-storage copy is part of the cluster's storage, so it depends on the cluster; whether any of it can be adopted by a fresh cluster is a platform-specific question you must answer before you rely on it, not during an incident.* ## What to do with the plan Move the line. Under *storage* write "long history is kept affordably off the local volumes". Under *continuity* write what actually answers a lost cluster — another cluster that is already running and fed. If someone insists the parked history is the backup, the cheap test settles it: stand up an empty cluster, point it at that storage, and see whether a reader can be served a single record without the original cluster's metadata. That rehearsal is worth more than the argument.

  • Which records are guaranteed not to be in the remote-storage copy when a cluster is lost?
    The newest ones. Only closed history is parked, so the active tail — the records written in the period right before the loss, which is what an incident is usually about — is still on the cluster's own volumes. Any plan that leans on the parked objects is implicitly accepting the loss of exactly the span people care most about.
  • A team says its parked history survives an accidental stream removal. How would you check?
    Test it rather than argue. On most platforms removal propagates to the parked objects because their lifecycle is the stream's, but this genuinely varies. Remove a throwaway stream on a non-production cluster and look at the storage afterwards. If the objects go, the claim is dead; if they remain, you still have to show a fresh cluster can adopt them.

saying these in an interview costs you the question

  • History sits in object storage, so the cluster can be rebuilt from it
  • Parked history survives deleting the stream, so it is a backup
  • The storage provider's durability means the stream cannot be lost
  • A reader can be pointed straight at the parked objects while the cluster is down
  • Parked history carries reader positions and permissions with it