skip to content

After a restart with no valid stored read position, a reader group silently skips a day of records, so which rule decided that and why was no error raised?

level: seniorimportance: should knowfreq 45%

answer

  1. no valid position, fallback applies
  2. newest skips, oldest re-reads
  3. out of range counts as invalid too
  4. a configured default is not a failure

basics

~20 s

The start-from rule ran. A reader with no valid position begins where that rule says, typically the newest record, the oldest retained record or a point in time, and beginning at the newest skips everything between. Applying a configured default is not an error.

solid answer

~50 s

A position stops being valid in two ways: it expired while the group was away, or it still exists but names a point before the oldest record still retained. In both cases the group has nothing usable to resume from, so the **start-from rule** decides the resume point instead. Where that rule says newest, the stretch written during the outage is passed over and the group looks perfectly healthy with nothing unread; where it says oldest retained, the group re-reads the whole retained window and does the work again. Neither is an error from the cluster's point of view — a configured default was applied exactly as configured — so on many platforms nothing is logged as a failure and nothing refuses to start. Platforms differ here: some fail closed and refuse to begin without an explicit choice, which turns a silent data hole into a loud startup failure.

go deeper

for a junior

Hold on to the shape: a reader with no usable record of where it got to has to start somewhere, and a configured rule decides that somewhere rather than the reader itself.

for a middle

Explain the two ways a position becomes unusable — it expired, or it names a point already removed — and what each of the common start-from choices costs when it fires.

for a senior

Demonstrate the diagnosis. Say which branch fired from the group's state right after start, say why nothing raised an error, and say what you would check downstream to size the damage.

for a principal

The judgment is about defaults across many groups: which groups may be fast-forwarded and which must never be, and whether you would rather a platform that fails closed and forces the choice.

## What "no valid position" actually means A reader group resumes from its **read position** only when it has one that the cluster will accept. There are several ways to arrive without one: - **It never had one.** A group name used for the first time — including a name that accidentally changed, because a renamed group is a new group with an empty history. - **It expired.** The **stored position lifetime** ran out while the group was inactive, even though the records are still inside the **retained window**. - **It is out of range.** The position survived but points before the oldest record still retained, because the records it named aged out while the reader was away. - **The stream it referred to was replaced.** On some platforms a stream recreated under the same name leaves earlier positions meaningless. All four land in the same state, and the state is what matters: the group cannot resume from its own history, so something else must choose. ## The rule that fills the gap That something else is the **start-from rule** — the policy for where a reader with no valid position begins. Three shapes are common, and they have very different costs: | Start-from rule | Where the group begins | What that costs | |---|---|---| | Newest | At the end of the stream, reading only what arrives from now on | Everything written during the outage is passed over, and no record of the hole exists | | Oldest retained | At the front of the retained window | The whole window is processed again, at full cost, with every downstream effect repeated | | A point in time | At the first record at or after a chosen timestamp | Only as good as the timestamp, which the operator has to know and get right | The rule is not a per-incident decision. It was chosen once, often as a default nobody revisited, and it fires unattended at exactly the moment a position turns out to be missing. ## Why nothing is raised as an error From the cluster's side, nothing went wrong. A group asked to start, it had no usable position, a configured default said where to begin, and the default was applied. That is ordinary operation, so on many platforms there is no failure to log, no refusal to start and no signal distinguishable from a genuinely new group beginning work. The information that something was lost lives only in the difference between where the group *would* have resumed and where it actually did — and by then the record of where it would have resumed is exactly what is gone. This is why the incident is usually discovered downstream: a gap in derived data, a report missing a day, an aggregate that is quietly wrong, or duplicated work if the rule sent the group backwards instead. ## Telling a skip from a re-read The two outcomes are opposite and immediately distinguishable once you know to look: 1. **Nothing unread, instantly.** The group reports itself fully caught up seconds after starting. That is the newest-record branch, and the outage window has been passed over. 2. **Unread equal to the whole retained window.** The group has an enormous amount to get through that nobody produced. That is the oldest-retained branch, and the work is being repeated. One clause of nuance: what you are looking at here is a symptom that identifies which branch fired, not a measurement exercise. The number's only job is to tell the two apart. ## Designs differ, and the difference is load-bearing - Platforms that **fail closed** refuse to start a group with no valid position and require an explicit choice. Noisier, safer, and the reason some teams prefer them. - Platforms that **default to newest** favour availability of the reading side over completeness, which is a reasonable default for live dashboards and a terrible one for ledgers. - Designs that **keep no stored position** have no start-from rule at all, because there is nothing to fall back from. An unacknowledged record simply remains until it is acknowledged or removed under whatever rule governs record lifetime, so this failure mode does not exist for them. ## Reducing the blast radius before it happens - Know your start-from rule by name for each group that matters, and know it before the outage rather than during it. - Make the rule's choice match the group's job: a group whose output must be complete should not be quietly fast-forwarded to the newest record. - Prefer a position lifetime long enough that the rule is rarely reached at all, which removes the question rather than answering it. - Treat the group name as part of the configuration. A name that varies with a host, a build or an environment produces a brand-new group with no position every time it changes, and the rule fires then too.

  • Both outcomes come from the same rule, so how do you tell which one actually fired?
    By the state of the group seconds after it starts. Fully caught up with nothing unread means it jumped to the newest record and the outage window was passed over. A backlog of unread records the size of the whole retained window means it went to the oldest retained record and is repeating work. Downstream tells the same story as a hole or as duplicates.
  • Besides expiry, what makes a stored read position invalid?
    Most often it points before the oldest record still retained, because the records it named aged out while the reader was away. A group name that changed is another: the new name has no history, so it is treated as a first start. On some platforms a stream recreated under the same name also leaves earlier positions meaningless.
  • Why can a group that skipped a day still look completely healthy?
    Because every operational signal it produces is about its current progress. It has a valid position, it is reading new records as they arrive, and it has nothing unread. Nothing in that picture encodes the records it never looked at, which is why the discovery usually comes from the data it was supposed to produce.

saying these in an interview costs you the question

  • Expects the broker to fail loudly whenever a position is missing
  • Thinks a reader always resumes at the oldest retained record
  • Believes a position pointing before the retained window is still honoured
  • Assumes the skipped records can be recovered from the position store
  • Treats a silent re-read as harmless because nothing reported an error
  • Ignores that a changed group name is also a group with no position