skip to content

A copied record gets a new position number on the target cluster, so what relates the two clusters' position spaces?

level: middleimportance: should knowfreq 48%

answer

  1. two clusters, two coordinate systems
  2. same record, different number
  3. only the copier sees both
  4. sampled pairs, written as it runs

basics

~20 s

Only a position map written by the cross-cluster copier while it ran: sampled pairs of a source position and the nearest position on the target. Without one, a number recorded against the source names no particular record on the target.

solid answer

~50 s

The stream on the target cluster is an independent append-only sequence, and it numbers what it writes in its own space. Even if both streams began empty and carry the same records, the numbers drift apart the moment anything differs - the hop started late, older records had already aged out of the source, a retry appended a duplicate, or the target stream already held something. So a number recorded against the source - a reader's stored read position, a job's checkpoint, an audit line saying 'processed up to here' - identifies no particular record on the target. The only thing that relates the two spaces is a **position map**: pairs of source position and the nearest target position, written by the copier as it goes. It is sampled, so it resolves to at-or-before a point rather than exactly, and it cannot be reconstructed afterwards - if the copier never wrote it, the correspondence is gone.

go deeper

for a junior

Recall the headline: the same record has a different number on each cluster, so a number copied from one side does not point anywhere useful on the other.

for a middle

Be able to explain why the numbers drift - a late start, records aged out, a duplicate from a retry - and to name the position map as the only thing relating the two spaces.

for a senior

The production judgment is to check, before you need it, whether the hop you run writes a map at all and how often it samples, because neither can be recovered after the source is gone.

for a principal

Frame it as evidence that has to be produced continuously at some cost, not a capability you can buy during an outage. That reframing is what gets the map turned on while the estate is healthy.

## Two independent numbering spaces On platforms that give each record a numeric position within a stream, that number is issued by the cluster doing the appending. The stream on the **target cluster** is a different stream on a different cluster: it numbers the records it receives in its own sequence, starting wherever its own sequence happens to be. The copier cannot ask the target to reuse the source's numbers, because the number is the cluster's own record of where it put something. So after a copy hop you hold two streams which may contain the same records and which have **two unrelated coordinate systems** over them. ## Why the numbers drift apart Even an optimist who assumes both streams started empty on the same day has to survive this list: - **The hop started later than the stream did.** Everything written before the copier existed is on the source only, so the target is that many numbers behind from record one. - **Records aged out before the hop read them.** Anything removed by the source's retention rule while the copier was stopped simply never crosses. - **The copier re-appends after an interruption.** A copy is at-least-once; a duplicate consumes a number on the target that corresponds to nothing new on the source. - **Filtering or reshaping.** If the hop carries a subset, or fans one source part across several target parts, the correspondence is not even one-to-one. - **The target stream was not empty.** A re-created hop into an existing stream starts at whatever number that stream had reached. Once any one of these has happened - and one always has - number 4,812,300 on the source and number 4,812,300 on the target are different records, and nothing about them warns you. ## What a stored source number is worth on the target Nothing, on its own. That includes every artefact that quietly holds one: | what holds a number | where it came from | what it means on the target | |---|---|---| | a reader group's stored read position | the source cluster | no particular record | | a downstream job's checkpoint | the source cluster | no particular record | | an audit row saying 'processed up to here' | the source cluster | no particular record | | a support ticket quoting a record number | the source cluster | no particular record | This is the deflating fact about mirroring that interviewers are looking for: the hop moves the records and leaves the coordinates behind. ## The position map The repair is a **position map**: a stored correspondence between a position on the source and the nearest position on the target, written by the copier *while it runs*, because only the copier is in a position to know both at once. Three properties matter: 1. **It is written continuously, not derived later.** Once the source cluster is gone, nothing can reconstruct which target record corresponds to a given source number - the evidence was in the copier's own hands at copy time. 2. **It is sampled.** The copier records a pair every so often rather than for every record, so resolving a source number gives you a target position at or slightly before the record you meant. In practice that means a little replay rather than a gap, which is the safe direction to err. 3. **It is per part.** Because numbering is per part of the stream, so is the correspondence; a single stream-wide number would not mean anything. What is done with the map later - re-establishing a reader group against the standby, choosing between replay and a gap at restart - is the failover conversation. The point that belongs to mirroring is narrower and comes first: **the map exists only if the copier wrote it**, and a hop configured without it has silently thrown away the ability to relate the two sides. ## Where platforms differ Not every platform poses this problem in the same form. - Where readers resume **by time** rather than by number, a timestamp can stand in for the map - but only if the copier preserved the source's record timestamps rather than stamping arrival, so the two questions are linked. - On **queue-shaped** platforms where a message is removed on acknowledgement and no rewindable number exists, there is nothing to map: what crosses is messages, and the notion of resuming at a number does not arise. - Some copiers write a map by default, some only when asked, and some do not offer one. That is a property of the hop you are actually running, and it is worth checking before you need it rather than during the switch.

  • Both streams started empty and the hop has never been interrupted. Can you assume the numbers match?
    Not safely. It may be true today and stop being true after one retry, one filtered record or one copier re-creation, and nothing tells you when it broke. Any procedure that depends on the numbers matching is relying on a coincidence that a duplicate append can end silently.
  • Can the position map be rebuilt after the source cluster is lost?
    No. The correspondence only ever existed inside the copier as it read from one side and wrote to the other. Once the source is gone you can still read the records on the target, but you cannot work out which target position a given source number referred to.
  • Why is the map sampled rather than recorded for every record?
    Recording a pair per record would roughly double the write volume of the hop for information almost none of which is ever read. Sampling at intervals costs little and resolves a source number to a target position at or just before the intended record, which errs toward replay rather than toward skipping records.

A courier photographing pages from a ledger into a fresh notebook. The words are faithful, but page 40 of the notebook is not page 40 of the ledger - and only the courier, while copying, can write down which page became which.

saying these in an interview costs you the question

  • Assumes a record keeps its number across the hop, so the two clusters' numbers line up
  • Believes the correspondence can be reconstructed after the source cluster is gone
  • Expects a reader's stored position from the source to mean something on the target
  • Thinks the position map is exact rather than sampled
  • Assumes every copier writes a position map whether or not it was asked to