Across an estate, how would you set stored position lifetime against the retained window, and what changes where readers keep the position themselves?
answer
- one relationship, not two numbers
- position should outlast its records
- add the longest planned outage
- own the store, own the lifetime
basics
~20 sMake the stored position lifetime at least as long as the retained window plus the longest reader outage you intend to survive unaided. Where the reading side keeps its own position, that store's cleanup and restore policy becomes the lifetime, and someone must own it.
solid answer
~50 sTreat the two as one decision made in two places. Take the retained window as given by whoever owns record lifetime, then set the position lifetime so a position cannot die before the records it points at — otherwise a long-enough outage quietly turns into a resume-from-a-rule event while all the data is still there. Add the longest outage you are willing to survive without manual repair: a holiday freeze, a paused environment, a service switched off during an incident. Where the position lives outside the broker, in a store the reading side owns, the lifetime is no longer a cluster setting at all; it is that store's expiry, cleanup, backup and redeploy behaviour, so the policy must name an owner for it. Where a design keeps no stored position, the equivalent question moves entirely to record lifetime and belongs to whoever owns that.
go deeper
The takeaway is that someone has to decide how long a reader's recorded place is kept, and that the decision only makes sense next to how long the records themselves are kept.
Be able to state the relationship rather than two numbers: a position that expires before the records it points at converts an ordinary outage into a resume-from-a-rule event.
Show the operational consequences you are protecting against by name, and show that you know the lifetime moves to a different owner entirely once the position lives in a store your service writes.
The call is about defaults, exceptions and ownership across many teams: one relationship written down, an exception list with reasons, and a named owner for every position store that is not the cluster's.
## One decision made in two places The **retained window** and the **stored position lifetime** are almost always configured by different people at different times: retention by whoever owns the stream and its cost, position lifetime by whoever set up the cluster and never looked again. Nothing in most platforms relates them, and nothing complains when they disagree. Yet their relationship is the entire risk: if a position can expire while the records it points at are still retained, then any reader outage longer than the position lifetime silently becomes a resume-from-a-rule event on a stream that still holds everything the reader wanted. So the estate-level move is not to tune either number in isolation. It is to make the pair a single policy with a single owner, stated as a relationship rather than as two independent values. ## A default worth writing down 1. **Take the retained window as given.** Whoever owns record lifetime owns it; this policy does not relitigate it. 2. **Set the position lifetime to at least that window.** A position that outlives the records it names is merely useless; a position that dies while its records live is a loss you did not need to take. 3. **Add the longest outage you intend to survive without manual repair.** Count honestly: an end-of-year freeze, a paused non-production environment, a reading side switched off deliberately during an incident, a migration that overran. 4. **Record the exception list.** Some groups genuinely should be fast-forwarded after a long absence — a live view whose old input is worthless. Make that an explicit exception rather than an accident of defaults. ## Three storage shapes, three owners | Shape | What the lifetime actually is | Who must own it | |---|---|---| | The broker holds the position per reader group | An inactivity timer configured on the cluster | The platform team, as a cluster-wide default with named exceptions | | The reading side writes the position into its own store | The store's row expiry, cleanup jobs, restore behaviour and whether a redeploy wipes it | The team that owns the reading service and that store | | No position is stored; records are removed on acknowledgement | Nothing, because there is no position; the unacknowledged records are the state | Whoever owns record lifetime, since that is the only lifetime present | The middle row is the one that quietly removes the platform team from the picture. Moving the position into your own store buys control, and it buys the whole lifecycle with it: the position is now application state that must be backed up, must survive a redeploy, and must not be cleaned by a job that was written for something else. Teams reach for it to escape a cluster default and then discover that an environment rebuild resets every reader to the start-from rule at once. The bottom row is not a loophole either. It removes the expiry question and replaces it with a harder one: with no stored position there is nothing to rewind to, so the authoritative answer to "what has not been handled" is whatever is still unacknowledged in the broker, governed by a lifetime this policy does not own. ## What the policy has to name - **The default relationship**, in words: position lifetime is at least the retained window plus the agreed outage allowance. - **The exceptions**, per group, with the reason recorded, so nobody has to reconstruct the intent later. - **The owner of each position store** that is not the broker, including its backup and its redeploy behaviour. - **The group naming rule**, because a name that varies by host, build or environment produces a brand-new group with no position every time it changes, which defeats any lifetime setting. - **A review trigger**: retention changes and reader-side platform changes both invalidate the pairing, and the pairing is what the policy is for. ## The failure the policy prevents The one worth describing to a reviewer is unremarkable and expensive: a reader is switched off deliberately, it stays off longer than anyone planned, its position ages out, it is switched back on during the recovery, and it resumes wherever a default decided. Nothing failed, nothing was logged as a failure, and the damage is a hole or a duplicate run in whatever the reader produced. Every element of it is a configuration relationship nobody owned. ## Where the question does not apply Be explicit about the boundary when you answer. On a design with no stored position the question dissolves, and on a hosted offering the lifetime may not be adjustable at all — in which case the policy is not a number you set but a constraint you plan around, by keeping readers from being absent that long or by accepting the start-from behaviour deliberately rather than by surprise.
- The platform is hosted and the position lifetime is not adjustable. What does the policy become?A constraint to design around rather than a value to set. You keep readers from being absent longer than the fixed lifetime, you decide the start-from behaviour deliberately for each group instead of inheriting it, and you record which groups would be materially damaged if the rule fired, so the risk is a known one rather than a discovery.
- Is there a cost to making the position lifetime very long?Some. The position store grows with every group name ever used, and abandoned groups stop being reclaimed automatically, so stale entries accumulate and have to be retired by someone. That is a governance chore rather than a systems risk, and it is a far cheaper problem than a reader that silently resumed from a default.
- Why is the group naming rule part of this policy at all?Because a lifetime protects only an existing position. If a group's name carries a host, a build number or an environment, every change produces a fresh group with no position, and the start-from rule fires regardless of how generous the lifetime was. Stable names are what make the lifetime setting mean anything.
saying these in an interview costs you the question
- Sets position lifetime and record lifetime independently and never compares them
- Thinks keeping the position in your own store removes the lifetime problem
- Assumes a hosted platform's position lifetime can always be raised
- Ignores that a changing group name defeats any position lifetime
- Believes a very long position lifetime has no housekeeping cost