The source cluster is lost and its uncopied tail dies with it, so what does that gap cost the systems downstream?
answer
- acknowledged upstream, absent downstream
- nothing on the target reports a hole
- replay cannot reach what never crossed
- cost depends on the system of record
- repair comes from the writer, not the stream
basics
~20 sThose records were acknowledged to their writers, so upstream believes the facts happened while every system derived from the stream is permanently short of them — and the gap is silent, because readers resuming on the target cluster see an unbroken stream with no hole in it.
solid answer
~50 sThe uncopied tail is not a delayed delivery; it is a set of records that existed on one cluster and now exist nowhere. Every writer was told those writes succeeded, so the upstream side of the business believes the facts are recorded. Everything derived from the stream — a ledger projection, a search index, a downstream service's own store — is missing them for good, and will stay missing until something outside the stream supplies them. The dangerous property is silence: readers pick up on the standby cluster and see a continuous stream, so nothing errors and no alert fires. The cost therefore depends entirely on whether the stream was the *only* record of those facts or a shadow of state held somewhere else, and that distinction — per stream, not per cluster — is what decides how big a loss window you can actually afford.
go deeper
Recall that the uncopied tail is gone rather than delayed, and that the writers were already told those writes succeeded. Upstream thinks the facts exist; everything downstream is missing them.
Explain why the loss is silent: readers resuming on the target cluster see an unbroken stream, and no component is responsible for noticing what never arrived. Say how you would size the gap afterwards.
Demonstrate that you price the loss per stream. Show which streams are repairable because a writer still holds the state, which are not, and why replay is never the repair for this.
Argue the posture: which streams may not be the system of record at all, what separating hops buys, and how to get an owning team to state on the record which loss they accept.
## The tail that dies with the source cluster The **uncopied tail** is the set of records that existed on the source cluster and had not reached the target cluster when the source was lost. It is tempting to think of it as data in flight that will turn up later. It will not. The copier reads from a cluster that is gone; there is no queue somewhere still draining. Those records have exactly one property that matters: they were **acknowledged to their writers**, and they now exist nowhere. That asymmetry is the whole cost. The producing side got a success and moved on — it marked an order placed, released a lock, returned a 200, deleted its own copy of the payload. The consuming side never sees the record at all. ## Why the gap is silent Nothing about a missing tail announces itself, and this is what turns a known loss window into an incident weeks later. - Readers that resume on the **target cluster** see a continuous stream. Where the stream is split into parts with their own position numbers, those numbers on the target are the target's own and run unbroken; where records are removed on acknowledgement, the missing items simply never appear. Either way there is no hole to trip over. - No component is responsible for noticing. The copier is gone with the cluster it was reading. The target cluster was never told what it did not receive. - Downstream systems are usually self-consistent afterwards. A projection built from a stream that is missing five thousand records is not corrupt; it is just wrong, and it looks fine. The practical consequence is that the loss surfaces as a reconciliation difference — a count that does not match, a customer who insists they placed an order — rather than as an alert. ## What it costs, stream by stream The stated loss window is the same for every stream on the hop, but what it *costs* is not. The useful question is whether the stream is the system of record for those facts or a derived view of something that survives elsewhere. | What the stream carries | What the tail's loss costs | Can it be repaired? | |---|---|---| | Facts whose only durable home is the stream (events captured at the edge, one-shot commands) | Permanent, unbounded — the fact is simply gone | No, except from whatever the writer still holds | | A shadow of state held in a writer's own store (change records emitted from a database) | Temporary divergence in every derived consumer | Yes — re-emit from the writer's store, or re-derive the consumer | | High-volume telemetry or metrics | A gap in charts and aggregates | Usually not worth repairing; the cost is accepted | | Instructions that cause external side effects (payments, notifications) | The side effect never happens while upstream believes it did | Only by reconciliation against the external system | ## Repairing what can be repaired There are exactly two honest repair routes, and neither one involves the stream: 1. **Re-emit from a source that survived.** If the writer holds its own durable state — a database whose change records became the stream, a request log — the tail can be regenerated. This is the single strongest argument for having writers keep state rather than treating the stream as the only ledger. 2. **Reconcile against the external world.** Where the records caused or recorded effects outside your systems, compare the downstream state to the external system of record and repair the difference by hand or by a targeted job. What does *not* work is replay. Replay reads from a stream, and the records were never on the surviving one. ## What this changes about the target you state This is why a stated recovery point is a per-stream decision and not an estate-wide one. A stream carrying change records from a durable store can genuinely afford thirty seconds of loss, because thirty seconds of divergence is repairable and the repair is understood. A stream that is the only home of a business fact cannot afford it in the same way, and the honest responses are narrow: - Give that stream a much tighter loss window, at whatever the copier capacity costs. - Move it to a synchronous cross-site write and pay the round trip on every write. - Give the writer a durable store of its own so the stream stops being the system of record — usually the cheapest of the three. - Or state the number plainly, in records, and get the owning team to say out loud that they accept it. An interviewer is listening for the last one as much as the first three. Naming the cost and getting it accepted is a real answer; assuming the loss is recoverable is not.
- Would a longer retention window on the target cluster have preserved any of the uncopied tail?No. Retention decides how long a record survives once it is on a cluster. The uncopied tail never reached the target, so there is nothing there for retention to keep. Retention protects against deleting history you have; it does nothing about history you never received.
- How would you detect after the fact that a tail was lost, and how big it was?Compare counts across the boundary: what writers believe they produced in the window against what arrived on the target cluster, using the writers' own records or an independent counter. The copy lag observed just before the loss gives an estimate of the span; the writers' side gives the actual figure. Neither is available from the stream alone.
- Does this change which streams belong on the same copy hop?Yes. Streams sharing a hop share its capacity and therefore its lag, so a high-volume stream with a cheap loss can push a low-volume stream with an expensive loss well past its stated ceiling. Separating them onto their own hop is often cheaper than tightening the target for both.
saying these in an interview costs you the question
- Expects the uncopied tail to arrive once the copier restarts
- Thinks replay on the target cluster can recover the lost records
- Assumes a gap raises an error for readers on the standby
- Sets one loss window for every stream regardless of what it carries
- Believes longer retention on the target would have preserved the tail
- Treats the stream as the system of record without saying so