skip to content

Stored Position Lifetime

How long a stored reader position outlives the reader itself, and what happens when it expires or points outside the retained window. Asked because the rule that then applies can silently skip a day.

part ofBroker & streaming operationsoverview, primer and where to startread it →
on this pageshow

questions

3

Why can a reader group's recorded read position vanish after an idle weekend while every record it had not read is still there?

level: middleimportance: must knowfreq 55%

answer

  1. two clocks, not one
  2. inactivity timer, not record age
  3. the position store has its own lifetime
  4. where it is kept decides what ends it

basics

~20 s

A stored read position and the records it points at live under two independent lifetimes. Many platforms discard the position after the owning reader group has been inactive long enough, while the records stay inside the retained window.

solid answer

~50 s

Records live inside the **retained window**; the read position lives in a position store that applies its own, usually much shorter, lifetime, and that lifetime is commonly measured from the last time the reader group reported progress rather than from the age of any record. So a group parked over a weekend, a frozen deploy or a holiday shutdown can come back to find the records intact and its place forgotten. Where the position is kept changes what ends it: a broker-held position is typically aged out by an inactivity timer, a position the reading side writes into its own store dies whenever that store is cleaned, restored or wiped, and a design that removes a record on acknowledgement keeps no position at all, so there is nothing to expire. Losing the position destroys no records; it only means the group resumes from a fallback rule instead of from where it actually got to.

go deeper

for a junior

Remember the plain fact: a reader group's recorded place and the records themselves are kept separately, so one can disappear while the other is still sitting in the stream.

for a middle

Be able to say what each clock measures — age or bytes of records for the retained window, and time since the group last reported progress for the position store — and that the two are set independently.

for a senior

Show that you plan for the mismatch. Name the outages long enough to outlive a position, such as a holiday freeze or a reading side deliberately switched off, and say how the return would be noticed.

for a principal

The angle is that these are one decision made in two places, often by two different owners, and that a position kept in a team's own store silently inherits that store's cleanup, backup and redeploy policy.

## Two lifetimes, not one Any platform that lets a reader resume where it left off keeps two different things alive. The first is the **records** in the stream, which live inside the **retained window** — the span the cluster still holds, whether that span is bounded by age, by bytes, or by a rule that keeps only the latest value per key. The second is the **read position**: a small piece of metadata recording how far one reader group has got. It is tiny, it is rewritten constantly, and it is stored somewhere else entirely — in the broker's own metadata, in a side store the reading side writes to, or, on designs that remove a record once it is acknowledged, nowhere at all. Those two things are governed by two clocks that have nothing to do with each other. The retained window is measured against records. The **stored position lifetime** is commonly measured against *inactivity*: how long it has been since the reader group last reported progress. Either clock can run out first, and the case that surprises people is the one where the shorter clock is the position's — the group's place is forgotten while every record it had not read is still sitting in the stream. ## Why a position is expired at all A platform that never discarded a position would accumulate one entry for every group name ever used, and that matters more than it sounds: - Reader group names are created casually — a one-off backfill, a laptop experiment, a test suite, a name that carries a build number or a host name. - The cluster cannot tell a retired group from a resting one. **Inactivity is the only proxy available**, so that is what it uses. - The metadata holding positions is often on the critical path for startup and recovery, so keeping it small keeps those fast. - Nothing reclaims an abandoned position otherwise, so timing it out is the default housekeeping. The expiry is therefore deliberate, not a bug. The defect is the **mismatch**: a lifetime short enough to catch a group that was only resting. ## Where the position lives decides what ends it | Where the position is kept | What ends its life | Who controls that | |---|---|---| | The broker's own metadata, per reader group | An inactivity timer running from the last reported progress | Whoever owns the cluster's settings | | A store the reading side writes to — a table, a file, an object | Row expiry, a cleanup job, a restore, or a redeploy that wipes the volume | The team that owns that store | | Nowhere: records are removed on acknowledgement | Nothing, because there is no position to expire | Not applicable; the unacknowledged records are the state | The third row is the one people forget when they move between platforms. If the design keeps no position, "my position expired" is not a failure it can have; what it has instead is records that wait until they are acknowledged or removed under a rule someone else owns. ## The two shapes of the mismatch 1. **The position dies first.** The group is down longer than the position lifetime but well inside the retained window. It returns to a stream full of the records it wanted and no memory of where it was. 2. **The records die first.** The position survives but now names a point before the oldest record still retained, because the stretch it pointed at aged out while the reader was away. The position exists and is useless. Both end in the same state — the group has **no valid position** — and something other than its own history then decides where it resumes. ## What an operator does about it - Establish which of the three storage shapes you are on before reasoning about anything else; the answer changes completely between them. - Compare the two lifetimes as numbers rather than as intentions. A position lifetime shorter than the retained window means any long-enough outage quietly becomes a resume-from-a-rule event. - Count the outages that run longer than people expect: a holiday freeze, a paused environment, an incident where the reading side was deliberately switched off, a migration that overran. - If the position lives in a store your team owns, put it on that store's backup and redeploy checklist. It is application state with a lifetime now, and nobody on the cluster side is watching it. - Keep the consequence straight: losing a position destroys nothing in the stream, other reader groups are unaffected, and the records are exactly where they were. What is lost is the answer to "where were we".

  • A reader group is running but its stream has produced nothing for a week. Is its stored read position at risk?
    It depends on what the clock actually watches. Where the lifetime runs from the last reported progress, a live group with nothing to report can look inactive and age out; where presence or a liveness signal refreshes it, it does not. Treat it as a question to answer for the platform in front of you rather than an assumption to carry between platforms.
  • Does removing a reader group's read position delete anything from the stream?
    No. The position is metadata about one group's progress. The records are governed by the retained window and by nothing that group does, and other reader groups keep their own positions and are unaffected. The only thing lost is that group's knowledge of where it was, which is why it then resumes from a fallback rule.

A library keeps a book on the shelf for a year but bins the slip recording which page you reached after two weeks with no visits. Come back in week three and the book is exactly where it was while your bookmark is gone, and a house rule, not your memory, decides whether you restart at page one or at the last page.

saying these in an interview costs you the question

  • Assumes a stored read position lives as long as the records do
  • Thinks lengthening retention automatically lengthens how long a position survives
  • Believes a vanished read position means records were lost
  • Expects the cluster to warn a group before its position ages out
  • Assumes every messaging design stores a read position at all
open as a page

After a restart with no valid stored read position, a reader group silently skips a day of records, so which rule decided that and why was no error raised?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The start-from rule ran. A reader with no valid position begins where that rule says, typically the newest record, the oldest retained record or a point in time, and beginning at the newest skips everything between. Applying a configured default is not an error.

open as a page

Across an estate, how would you set stored position lifetime against the retained window, and what changes where readers keep the position themselves?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Make the stored position lifetime at least as long as the retained window plus the longest reader outage you intend to survive unaided. Where the reading side keeps its own position, that store's cleanup and restore policy becomes the lifetime, and someone must own it.

open as a page