skip to content

questions

3

Two sites accept writes to a stream of the same name, with a cross-cluster copier running each way — what stops a record circulating forever?

level: middleimportance: must knowfreq 52%

answer

  1. symmetry is the defect
  2. a copier is just another writer
  3. each lap re-copies the copy
  4. mark where the record first landed
  5. stamp, or origin-qualified stream name

basics

~20 s

An origin stamp: a marker on the record, or an origin-qualified stream name, that tells the copier running the other way this record came from elsewhere. Without one, each copier treats the other's output as new local traffic and the loop never ends.

solid answer

~50 s

A cross-cluster copier is just another writer on its target cluster — it appends records the target did not produce locally, and an append carries no inherent `came from elsewhere` flag. So with a hop each way, the copy that lands at site B is read by the B-to-A copier as ordinary new work and sent back, where the A-to-B copier picks it up again. The loop is self-feeding: every lap appends one more record at both sites, so it grows rather than settling. The fix is an origin stamp — either a marker carried beside the payload recording the site that first accepted the record, or writing copies into an origin-qualified stream name — so each copier forwards only what originated on its own side. Both shapes exist because platforms differ in what metadata survives a hop.

go deeper

for a junior

Recall that copying both ways with nothing marking where a record started sends the same record back and forth. Being able to name the echo loop and say it does not stop by itself is enough at this level.

for a middle

Explain that the copier appends exactly like a local writer, so the return hop cannot tell a copy from fresh traffic, and describe both shapes of the origin stamp: a marker beside the payload, or an origin-qualified stream name.

for a senior

Show that you would check the stamping arrangement before any second hop is enabled, and that you know the echo grows without bound rather than settling, spending retention, disk and copy lag at both sites while readers see the same payload repeatedly.

for a principal

Argue whether the estate should run two-way copying at all. Stamping stops the echo but leaves readers absorbing duplicates and two interleavings, so the discipline has to be maintained on every future hop by every future operator.

## The arrangement being described An **active/active** pair is two **sites** — a site being one cluster plus the writers and readers attached to it — where writers at *both* sites append to a stream carrying the same name, and the operators want each site to end up holding everything. The obvious first build is symmetric: one **copy hop** from site A to site B, and a second hop from site B back to site A, each run by a **cross-cluster copier**, the process that reads records from a source cluster and appends them to a stream on a target cluster. That symmetry is the defect, and the reason fits in one sentence: **a copier is just another writer on its target cluster.** It appends records the target did not produce locally, and an append carries no inherent 'this one came from somewhere else' flag. To the copier running the other way, those freshly arrived copies are simply records on its source cluster that it has not forwarded yet. ## The loop, one lap at a time 1. A producer at site A appends record `R` to A's stream. 2. The A-to-B copier reads `R` and appends it to B's stream. 3. The B-to-A copier, reading B's stream, sees that copy as new work and appends it to A's stream — a second `R` on A, at a later position on A. 4. The A-to-B copier sees that second `R` and forwards it to B, and the cycle repeats. No step in that list terminates. Nothing compares payloads, and nothing on either cluster treats 'we already hold something like this' as a reason to refuse an append. Each lap costs one appended record **at both sites**, so the arrangement never settles into a steady state — it grows. What ends it in practice is always something else breaking: the retention rule expiring records faster than the loop creates them, a volume filling, copy lag between the two clusters climbing because both copiers are spending their throughput on their own echo, or an operator noticing readers receive the same payload over and over. The property worth being able to state out loud is that the echo is **unbounded and self-feeding**, not a one-time duplication. A candidate who answers 'the record gets copied twice' has not seen the shape. ## The origin stamp The fix is to make a copied record identifiable *as* copied, so the copier running the other way can decline to send it back. This tree calls that marker an **origin stamp**, and it turns up in two shapes, because platforms differ in what they can carry across a hop. | Shape | Where it lives | How the return copier uses it | What it demands | |---|---|---|---| | A marker beside the payload | a metadata field attached to each record | skip any record whose recorded origin is not this source site | the platform must carry per-record metadata through the hop | | An origin-qualified stream name | the name of the stream the copy is written into | copy only the streams this site originates, never one already naming another origin | readers consume more than one stream, and stream names become load-bearing | Both encode the same fact: the record, or the stream holding it, carries the identity of the site that first accepted it, which turns 'only forward what originated here' into a checkable rule. The first shape keeps one stream name everywhere and hides the plumbing. The second needs no per-record metadata at all and is visible in a plain stream listing, which is why it survives on platforms whose records carry little beside the payload. ## What the stamp fixes, and what it does not - **Fixed:** the unbounded echo. A record makes one hop per direction and stops. - **Not fixed:** duplication in general. A copier restarted from an earlier point re-appends records it had already carried, and readers absorb those. - **Not fixed:** order. Records accepted independently at the two sites arrive interleaved by copier timing, and that interleaving is not the same at both sites. - **Not fixed:** the bookkeeping around records. Stored read positions, access rules, stream settings and payload contracts do not ride along on a copy hop. - **Not fixed:** the premise. Stamping makes writing at both sites *survivable*, not *correct*. ## Where platforms genuinely differ The loop can occur wherever two copiers face each other, but it presents differently: - On a platform that splits a stream into parts and gives a reader a rewindable numeric position, the loop shows up as relentless growth in record count and stored bytes, and a copy is distinguishable from local traffic only by its metadata. - On a queue-shaped broker where reading removes the record, the same loop presents as a queue that never drains: each delivery to the copier takes one record away here and creates one there, indefinitely. - Some platforms attach per-record metadata that survives a hop and some do not — which is exactly why the origin-qualified stream name exists as a real alternative rather than as a quaint older style. The durable habit is to describe what the copier *does* — reads from one cluster, appends to another, with or without a way to tell a copy apart — rather than to reach for one platform's tooling.

  • Why does the echo not settle once both sites hold the record?
    Because nothing anywhere compares payloads. Each cluster's stream is append-only from the copier's point of view: the returning copy is a new record at a new position, so the outbound copier sees new work and forwards it again. Every lap adds one record at each site, so volume grows until retention, a full volume or an operator ends it.
  • A copier restarts and loses whatever it knew about what it had already forwarded. Does the origin stamp still protect you?
    Against the echo, yes — the stamp travels with the record or with the stream name, so a restarted copier still recognises foreign-origin records and skips them. What a restart costs you is duplicates: resuming from an earlier point re-appends records already carried, and readers absorb those. Those re-appends are still stamped, so they do not loop.
  • What does an origin-qualified stream name cost the reader side?
    A reader that consumed one stream now consumes one per origin, and gains another whenever a site is added. The loop prevention has been moved into the name, so the fan-in moves into each reader's subscription list. That is a real cost, but it is visible and local, which is why platforms with thin record metadata prefer it.

Two office mailboxes with automatic out-of-office replies pointed at each other. The first reply triggers the second, which triggers the first, and neither program ever asks whether it has seen this exchange before. Only a rule of the form 'never auto-reply to something that is itself an auto-reply' — a marker saying where the message came from — ends it.

saying these in an interview costs you the question

  • Says the loop stops on its own once both sites hold the record
  • Believes the echo duplicates each record exactly once
  • Assumes a copier automatically knows which records it wrote itself
  • Calls the loop a network fault and proposes retry limits
  • Thinks more copies inside each cluster would have prevented it
  • Claims consumer-side duplicate handling makes the arrangement safe
open as a page

Both sites accepted writes to a stream of the same name during an hour-long network split. Why can no merge afterwards recover the true order across them?

level: seniorimportance: should knowfreq 45%

basics

~20 s

There was never one order to recover: each site sequenced only its own writes, and nothing anywhere recorded how the two interleaved. Any merge has to invent a rule — write timestamps, site priority — and different rules produce different, equally unfounded answers.

open as a page

A team wants one shared stream writable at both sites — why do operators propose one stream per site instead, and what does that cost readers?

level: principalimportance: should knowfreq 38%

basics

~20 s

One stream written at two sites has no defined order across them and needs loop prevention on every hop just to exist. Giving each site its own stream with one owning writer, read everywhere, removes both problems and moves the combining work onto readers.

open as a page