A team turns on keyed retention to shrink a stream before a deadline — what does keeping only the latest value per key actually promise?
answer
- a promise about survival, not removal
- no deadline attached to anything
- latest value per key, guaranteed readable
- footprint follows key count
- an age bound is the size lever
basics
~20 sKeyed retention promises only that the latest value for every key stays readable. It never states when a superseded value is removed, so it is a guarantee about what survives, not a size or deadline lever.
solid answer
~50 sKeyed retention — keeping only the latest value per key — guarantees what survives, not what goes. On platforms that offer it, a reader starting at the earliest position still on the store and reading to the end ends up holding the most recent value for every key that has one. Nothing states that the older, superseded values for those keys are gone by any particular moment: they merely become eligible for a background removal pass that runs when the platform decides, never touches the active segment, and can fall behind. The footprint of such a stream is therefore set by how many distinct keys exist and how large a value is, not by a number you configured. If a store has to fit a volume by a date, the instruments for that are an age bound and a byte ceiling; keyed retention is a modelling decision about the stream's contract.
go deeper
Recall the one-line promise: with keyed retention the store keeps the latest value for each key and may discard the older ones. Notice that 'may' is doing real work — nothing says when.
Explain the asymmetry. The rule is a guarantee about what survives and is silent about timing, because removal is done by a background pass over closed segments that never touches the segment still being appended to.
Show the operating consequence: you cannot commit to a footprint or a date with this rule. Size a keyed stream from distinct key count, value size and copies, and add an age bound when a number has to be met.
Frame it as a contract decision. Keyed retention trades the replay budget for superseded values against a bounded path to current state, and it is hard to reverse — turning it on destroys history that turning it off will not bring back.
## Two rules, two different promises A store holding a stream can be told to remove records in one of two ways. **Age-based removal** (oldest-first) says anything older than a bound may go. **Keyed retention** — keeping only the latest value per key — says something structurally different: any record for which a newer record carrying the same key exists may go. Not every platform in this class offers the second rule. Designs that remove a record once a subscriber acknowledges it have no notion of it, and some rented brokers expose only the first. Where it is offered, the guarantee runs in exactly one direction: **for every key that has been written and not explicitly deleted, the most recent value stays readable**. A reader that begins at the earliest position still on the store and reads to the end finishes holding the current value of every such key. That is the whole contract, and it is a property about survival. Three things it pointedly does not promise: - **When** a superseded value disappears. It becomes *eligible* for removal, and nothing states a deadline. - That a superseded value is unreadable in the meantime. A reader crossing old history may cross several earlier values for the same key before reaching the current one. - That the stream reaches, or stays under, any particular size. ## Why it is not a capacity lever The removal that keyed retention authorises is carried out by a background **removal pass**. The pass reads closed segments, keeps the surviving record for each key, writes the result into new segments and then releases the old ones. It never touches the **active segment** — the one still being appended to — so the newest slice of the stream always holds every value written into it, superseded or not. Platforms differ in how they decide when to run it: some schedule by elapsed time, some by how much of the eligible region is worth rewriting, and some expose no control at all. None of them offers what a capacity plan needs, which is *this stream will be under a stated number of bytes by Friday*. | What an operator wants to know | Keyed retention | Age-based removal | |---|---|---| | What is guaranteed to survive? | the latest value of every live key | everything newer than the bound | | What is guaranteed to be gone, and by when? | nothing, at no stated time | everything past the bound, once the pass runs | | Is the resulting size predictable? | only from key count and value size | roughly, from write rate times the bound | | Does it bound how far back history goes? | no | yes | ## What actually sets the floor A keyed stream's floor is **the number of distinct keys, times the size of a value, times the number of copies the cluster keeps** — plus the superseded bytes the pass has not reclaimed yet, plus the active segment. That has consequences worth saying out loud: - If keys are never reused — a fresh identifier on every record — nothing is ever superseded, and keyed retention reclaims nothing however long you wait. - If values are large, the floor is large even at modest key counts, because every live key contributes one value. - If the write rate is high relative to how fast the pass rewrites, the unreclaimed remainder is not a rounding error and can dominate the total. - Switching the rule on moves nothing by itself: no bytes are released until a pass has run over the closed segments. ## When the deadline is real If a store has to fit a number by a date, the instruments are **an age bound** (the maximum age a record may reach before it becomes removable) and **a byte ceiling** (the maximum bytes a stated scope may occupy), because each is enforced against a quantity you chose. Keyed retention answers a different question: what must a reader be able to reconstruct, and in how much work. The two questions are independent, which is why some platforms let both rules apply to one stream — and that combination removes more than people expect, because a record past the age bound goes even when it is the latest value for its key. It is worth being explicit about what is being traded. The retention window — the span of history still readable on the store right now — **is the replay budget**: it is the amount of history anything downstream can be rebuilt from. Keyed retention deliberately gives that budget up for superseded values, in exchange for a bounded path to current state. Choosing it is a statement about the stream's contract with its readers, not a knob to reach for when space is short.
- Does a reader replaying a keyed stream from the earliest position still on the store see each key exactly once?No. Until the removal pass has rewritten the closed segments, several superseded values for the same key can still be present, in write order, and the active segment always holds every value written into it. A reader that keeps its own copy of the latest value must simply let a later value overwrite an earlier one; the ordering within the key is what makes that safe.
- A stream under keyed retention uses a fresh identifier as the key on every record. What does the rule do for it?Nothing. No record is ever superseded, so the removal pass has nothing to reclaim and the footprint grows with every write. Such a stream needs an age bound or a byte ceiling, or a key scheme where identities actually repeat; keyed retention only pays where the same key is rewritten.
- Is keyed retention available everywhere in this product class?No, and the answer should say so rather than assume. It belongs to stores that keep a readable history addressed by key. A design that removes a record when a subscriber acknowledges it has nothing to keep the latest value of, and a rented broker may simply not expose the choice — in which case bounding by age or bytes is the only instrument available.
saying these in an interview costs you the question
- Says keyed retention shrinks a stream to a size you can predict from a setting
- Claims a superseded value is gone the moment a newer one is written
- Reaches for keyed retention instead of an age bound when a volume must fit
- Assumes every platform in this class offers the rule at all
- Expects a replaying reader to meet each key exactly once
- Thinks enabling the rule reclaims bytes immediately, with no pass involved