skip to content

An erasure obligation is executed against the primary stream alone, so which copies outside this cluster still hold the same bytes?

level: seniorimportance: should knowfreq 44%

answer

  1. scoped by bytes, not by cluster
  2. mirror, standby, offload, backup, sink
  3. an immutable image cannot be edited
  4. one key destruction reaches every ciphertext
  5. report the residue, not just the act

basics

~20 s

A mirrored cluster, a standby, segments offloaded to a remote object store, every backup image taken while the records existed, and every downstream sink fed from the stream. An erasure scoped to one cluster is not an erasure, because the obligation follows the bytes.

solid answer

~50 s

The obligation is scoped by **the copies of the record**, not by the cluster you happen to be logged into. The same bytes normally exist in at least five other places: an asynchronously mirrored cluster, a standby kept for failover, closed segments offloaded to a remote object store, point-in-time backup images taken at any moment while the records were present, and whatever downstream sinks were fed from the stream. Each has its own removal rules, and backups are the worst: they are immutable images that cannot be edited, so the only honest options are to wait for their own cycle or to hold the plaintext hostage to something you can destroy. That asymmetry is the strongest practical argument for per-subject encryption: destroying an encryption key reaches every copy holding that ciphertext at once, whereas a record-by-record erasure must be executed separately, and proven, in each place.

go deeper

for a junior

Recall that the same records exist in more than one place - a mirror, a backup, a remote object store - and that erasing one place is not erasing the data.

for a middle

Explain why each copy has its own removal rules, and why an immutable point-in-time image admits no selective edit at all.

for a senior

Show that you enumerate before you execute, rank the copies by difficulty, and close by stating exactly what residue remains after the work.

for a principal

Own the estate view: the longest backup lifetime and the set of sinks are numbers the organisation must know before it can promise an erasure deadline to anyone.

## Erasure is scoped by the bytes, not by the cluster The most common failure in this area is not technical, it is a scoping error. Someone executes an erasure against the stream they administer, records it as done, and never asks where else those bytes went. The obligation attaches to a subject's data wherever it is readable, and a broker cluster in production is a machine for making copies. ## The inventory At minimum, ask about all of these before calling an erasure complete: - **A mirrored cluster.** An asynchronous copy of the stream running elsewhere. It has its own bounds, and they are frequently more generous than the source's because the second cluster exists to be a fallback. - **A standby.** The cluster held ready for failover. It may be the same thing as the mirror, or a separate lineage with a separate operator. - **Offloaded segments.** Closed segments moved to a remote object store so history is no longer bounded by the local data volume. These outlive the local copy by design, and their removal is governed by the remote store's own rules. - **Backup images.** Point-in-time copies taken while the records existed. Each is a frozen picture of the store, and they are usually the longest-lived copy of all. - **Downstream sinks.** Anything fed from the stream that persists what it read. Each is a separate system with a separate owner and, usually, no record of which stream it came from. - **A reader that materialised it.** A service that rebuilt state from the history now holds a derived form of the same facts in its own store. | Copy | Who controls its removal | Why it is easy to miss | |---|---|---| | Mirrored cluster | The operator of the second cluster | Often a different team and a different bound | | Offloaded segments | The remote object store's own rules | Not visible on the local data volume at all | | Backup images | The backup cycle, and nothing else | Immutable by design; no selective edit exists | | Downstream sinks | The sink's owner | Nobody tracks the lineage backwards | ## Why backups are the hard case A mirror can be erased the way the primary was. A sink can be asked. A backup image cannot be edited at all — that is the property that makes it a backup. So there are only three honest positions, and you should be able to state which one you are taking: 1. **Wait for the cycle.** The image ages out under its own schedule, and the obligation's owner accepts a bounded delay during which the data remains restorable but is not in active use. This is the common negotiated answer, and it requires knowing the longest backup lifetime in the estate as a number. 2. **Re-erase on restore.** The erasure list is re-applied as a mandatory step whenever an image is restored, so the data can exist in an image but can never come back into service. This is only credible if the restore procedure genuinely cannot be run without it. 3. **Make the image useless for that subject.** If the payload was encrypted under a per-subject encryption key, destroying that key covers the backup as thoroughly as it covers the live store, because every copy holds the same ciphertext and only that ciphertext. The third is why an inventory exercise so often ends in an encryption decision. A record-level erasure has to be executed separately in every place and **proven** in every place; key destruction is one act whose reach follows the ciphertext, and it reaches every copy holding that ciphertext, and only those. It still does not reach a downstream sink that decrypted the payload and stored the plaintext, which is exactly the boundary worth naming out loud. ## Naming the residue honestly Whichever route you take, finish by saying what remains. After a record-level erasure on the primary, the derived aggregates a sink already computed still exist. After key destruction, the ciphertext, the record counts, the timings and any unencrypted routing field still exist. An engineer who reports "erased" without that sentence is reporting something they did not verify; an engineer who reports "erased on the primary, the mirror and the offloaded segments, with the longest backup image ageing out in thirty-one days, and these two sinks confirmed" is reporting the truth. ## What varies between platforms Do not assume the copies are the same everywhere. Where a platform keeps a replayable history, the stream itself is usually the longest-lived copy and the inventory starts there. Where a broker discards each record once it has been acknowledged, the store holds it for minutes and every surviving copy is downstream, so the same inventory has a completely different centre of gravity. Where the cluster is rented, the mirror, the offload target and the backup schedule may all be the provider's and may not be individually addressable at all — in which case the inventory question becomes a contract question, and the answer is whatever the provider will commit to in writing. ## What an interviewer is listening for That your first move is to enumerate rather than to execute. The named copies matter, the ranking by difficulty matters more, and the closing sentence that states precisely what still exists is what separates a candidate who has run one of these from a candidate who has designed one on a whiteboard.

  • Why is a backup image the hardest copy to reconcile with an erasure obligation?
    Because immutability is the property that makes it a backup — there is no selective edit, and rewriting it would destroy its value as evidence of a point in time. That leaves waiting for its own cycle, re-applying the erasure as a mandatory restore step, or having encrypted the payload so that destroying a key covers the image too.
  • Does destroying a per-subject encryption key cover a downstream sink?
    Only if the sink stored the ciphertext. A sink that decrypted the payload and persisted the plaintext, or computed something from it, is holding data the key no longer governs. Those sinks stay on the erasure list as separate systems with separate owners, and they are the boundary worth naming explicitly when you report the work as done.

saying these in an interview costs you the question

  • It is erased because it is gone from the primary cluster
  • The mirror will catch up and erase it automatically
  • Backups do not count because they are not in production
  • Offloaded segments follow the local data volume's bounds
  • Downstream sinks are the other team's compliance problem
  • Key destruction covers a sink that stored the decrypted payload