skip to content

Restarting a reader group on the standby means either reprocessing records it already handled or skipping some — how do you choose?

level: seniorimportance: should knowfreq 46%

answer

  1. no restart lands exactly where you were
  2. two errors: handled twice, or never
  3. silent is worse than visible
  4. overlap wider than the uncertainty
  5. the exception: only the newest matters

basics

~20 s

Choose per stream, by what a repeat and a miss each cost downstream. Default to landing early and reprocessing, since a duplicate is usually recoverable and an unread record is not, then make the overlap wider than the restart's uncertainty.

solid answer

~40 s

Because neither a position map nor a timestamp lands exactly where the group was, the restart is deliberately biased in one direction. Ask what each error costs for that stream: reprocessing means a record is handled twice, which is survivable where the handler tolerates repeats and unacceptable where it sends money or messages to a customer; skipping means a record is never handled at all, which is usually worse and always silent. The common default is to land early by a margin larger than the uncertainty — minutes, not seconds — so the overlap definitely covers the gap. The exceptions are streams where staleness is worse than absence, such as a feed of current values where only the newest record matters. Decide and record this per stream before the incident.

go deeper

for a junior

Recall that a reader restarted on another cluster lands near its old place, not on it, so it either re-reads a few records or misses a few — and that this is a choice, not an accident.

for a middle

Explain both costs and why they are asymmetric: reprocessing is usually visible and recoverable, skipping is silent and often permanent. Say why the margin should exceed the restart's uncertainty.

for a senior

Show that you decide per stream against downstream consequence, that you size the overlap deliberately, and that you have a third answer for consumers with one-way external effects.

for a principal

Make the per-stream decision an owned artefact recorded before the incident, with the overlap stated as a number, so the switch is executed from a prepared list rather than argued during an outage.

## Why there is a choice at all A **reader group** restarted on a **standby cluster** cannot land exactly where it was: a **position map** gives the nearest entry at or before the group's old position, and a timestamp restart gives the first record at or after a chosen moment. So the group either begins slightly behind where it was, and re-reads, or slightly ahead, and never sees the records in between. There is no third option, and pretending the restart is exact is how a switch becomes an invisible data-loss event. The decision is therefore: **which error do I want, and how much of it**. ## Frame it as two costs, per stream | | reprocessing (land early) | skipping (land late) | |---|---|---| | what happens | some records are handled twice | some records are never handled | | visibility | usually visible downstream | silent unless something reconciles | | typical harm | duplicated side effects | missing state, missing money, missing notification | | recoverable later | often, by reconciliation | only if the records still exist somewhere | | cost to the switch | extends the restart | none — it is faster | Two things fall out of that table. First, skipping is the more dangerous default precisely because it is cheap and quiet: the group catches up instantly and nobody sees anything. Second, the cost of reprocessing lands on the handler, and whether a repeated record is harmless is a property of the consuming application, not of the broker — that is assumed knowledge here, and it is the thing you check before promising that replay is safe. ## The rule that usually holds For most streams: **land early, and land early by more than you think you need**. The uncertainty in a restart is not one record; it is however coarse your map is, plus however long the group had been running without storing its progress, plus clock differences if you used a timestamp. An overlap measured in minutes costs the switch some extra reprocessing time and removes the entire class of silent gaps. The streams that genuinely invert this rule share a shape: **only the newest record matters**. A feed of current prices, positions, or device readings gains nothing from replaying the past ten minutes, and replaying it delays the consumer from reaching the present — which is the only thing it is for. For those, restarting at the newest record on the standby is right, and the skipped records are not a loss because nobody wanted them. A third case is worth naming: streams whose consumers produce **externally visible one-way effects** — sending a message to a customer, moving money, calling a third party. Here a repeat is not merely untidy, and the honest answer is often neither replay nor gap but a deliberate pause: restart the group with its effects suppressed, let it work through the uncertain span, and re-enable them once it is past the overlap. ## How to make either choice survivable 1. **Record the choice per stream before the incident**, next to the stream's owner. During a switch you want to read the decision, not make it. 2. **Prefer a margin you can state.** "Start ten minutes early" is auditable; "start where the map says" hides how much overlap you actually took. 3. **Make the overlap measurable afterwards.** Knowing how many records were re-read turns a vague worry into a number someone can reconcile against. 4. **Never let one blanket rule cover every stream.** A single global restart policy will be wrong for the feed that only wants the present and wrong for the stream that moves money, and both will be discovered afterwards. ## Where platforms differ - Where the reader owns a rewindable position, both directions are available and the choice above is real. - On **destructive-read** designs there is nothing to rewind: acknowledged messages are gone, and what you hold is whatever the copier carried and nobody consumed. The choice collapses into a different one — whether to re-drive the work from its original source or accept the absence. - Where the records are retained only briefly on the receiving cluster, "land early" has a floor: you cannot rewind past the oldest record the standby still holds, which can quietly turn an intended replay into a gap. - Some platforms let a group's progress be set administratively while it is stopped, and others require the group to start and seek; the mechanics differ but the decision does not.

  • Why is skipping treated as the more dangerous direction even though both are errors?
    Because it is silent and it is cheap. A group that lands late catches up immediately and looks healthy, and nothing in the stream says which records were never read. Reprocessing, by contrast, usually announces itself downstream — duplicated rows, repeated effects, a reconciliation mismatch — so it gets found and fixed.
  • Which streams genuinely justify restarting at the newest record instead?
    Those where only the present has value: current prices, live positions, device readings, anything where a consumer holds the latest value per key and older records add nothing. Replaying them delays the consumer from reaching the present, which is the one thing it exists to do, so the skipped span is not a loss.
  • What do you do for a consumer that sends customer-visible messages or moves money?
    Do not choose between a repeat and a miss for it. Restart the group early, with its outbound effects suppressed, let it work through the uncertain overlap, then re-enable them once it is past. You pay a controlled delay instead of either duplicate side effects or an unnoticed omission.

It is like resuming a book in a second copy with different page numbers. You either back up a few pages and re-read a little, or turn forward and hope you missed nothing — and since you cannot tell afterwards which pages you skipped, re-reading is the cheaper mistake unless you only ever wanted the last page.

saying these in an interview costs you the question

  • Applies one restart rule to every stream on the cluster
  • Assumes landing late is safe because the group catches up quickly
  • Chooses the overlap during the incident rather than beforehand
  • Treats the restart as exact and denies either error occurs
  • Promises replay is safe without checking what the handler does twice
  • Forgets the standby may not retain enough history to land early