How would you choose an age bound and a byte ceiling for a stream that several teams read, and what number should drive the choice?
answer
- size it backwards, not forwards
- Friday evening to Monday morning
- peak rate, never the average
- the ceiling is a guard, not a policy
- who pays for the extra span
basics
~20 sDerive the age bound backwards from the longest a reader may be down and still catch up — outage plus detection plus repair — then size the byte ceiling from the peak write rate so it guards the data volume without ever binding first.
solid answer
~50 sStart from a recovery requirement, never from a round number of days. Ask how long a reader may be stopped and still catch up: the outage itself, plus the time to notice it, plus the time to fix and redeploy — a Friday-evening failure noticed on Monday is already three days. That span, plus headroom, is the **age bound**. Then convert it to bytes at the **peak** write rate, not the average, and set the **byte ceiling** above that figure per whatever scope the platform counts, so the ceiling only ever fires as a guard on the data volume. Two supporting decisions make it stick: an estate-wide default so most streams need no thought, and an explicit owner for any stream that asks for more, because a longer retention window is a bill somebody pays.
go deeper
The takeaway is the question, not the arithmetic: retention exists so a reader that stops can still catch up, so the span should come from how long it might be stopped.
Be able to convert a chosen span into bytes at a stated write rate, and to say which scope that byte figure is counted against before calling it a capacity number.
Argue the headroom: peak rather than average rate, plus a factor for uneven routing, so the ceiling never binds during the very traffic that makes history valuable.
Own the estate policy — one sane default, an exception route with a named owner and a visible cost, and a review trigger for streams whose traffic has outgrown their bounds.
## Start from the outage you must survive The only defensible way to choose an age bound is backwards, from a question about people and process rather than about storage: **how long may a reader be down before the records it has not yet read are gone?** That span is almost always underestimated, because it is a sum: 1. the outage itself — a bad deploy, a dependency down, a cluster lost; 2. the time to **notice** — which on a weekend or a holiday is the largest term by far; 3. the time to diagnose, fix, review and redeploy; 4. the time to **catch up** afterwards, during which the front of the history is still moving away. A failure at 18:00 on a Friday, noticed at 09:00 on Monday, fixed by that afternoon, is already three days before catch-up begins. A one-day retention window would have destroyed the records days before anyone looked. This is why "three days" and "seven days" are the two spans in common use: they are Friday-to-Monday plus a margin. ## Turn the span into bytes Once the span is chosen, the byte ceiling is arithmetic — and it must use the **peak** rate, because the average is exactly the number that made the ceiling bind during a surge: ``` recovery span wanted = 5 days (3 days outage window + 2 days margin) peak write rate = 90 GB/day for this stream bytes for the span = 5 x 90 = 450 GB scope = per partition of a stream, 24 partitions per-scope share = 450 / 24 = 18.75 GB byte ceiling per scope = 18.75 x 2 (skew + surge headroom) = ~38 GB worst-case stored history = 38 x 24 = ~900 GB, which the data volume must hold ``` The doubling is not decoration. Records are rarely routed evenly, so the busiest partition reaches its ceiling long before the average one would, and without headroom the ceiling binds on the hot partitions during exactly the traffic that made you want the history. ## What each bound is for | | role | how you size it | it binding means | |---|---|---|---| | age bound | the promise made to readers | backwards from the recovery span | working as intended | | byte ceiling | the guard on the data volume | forwards from the peak write rate | the promise has quietly lapsed | That second row is the whole design rule: **a byte ceiling that binds in normal operation has silently replaced your retention policy with whatever today's traffic allows.** Size it so it does not, and monitor the readable span so you learn on the day it starts to. ## The estate view A lead is not setting one stream, they are setting a policy for hundreds: - **A default that fits most streams**, so that the great majority are created with no conversation at all. Uniformity is worth real money here: a bespoke span per stream produces an estate nobody can reason about or cost. - **An exception route with an owner.** A team asking for a much longer retention window is asking for capacity, and the request should name who is accountable for it and what recovery it buys. "In case we ever need it" is not a recovery requirement. - **A cost signal that reaches the asker.** Storage for a long history is invisible to the team that requested it unless someone makes it visible, and invisible costs never shrink. - **A review trigger.** Streams whose write rate has grown several-fold since their retention was set are the ones whose promise has already lapsed. ## The line to hold There is a strong pull towards "keep everything" — history feels free until it is not, and every bad week produces a request to extend. The honest framing to bring to that conversation: the retention window **is the replay budget**, it is bought with storage, and a team that wants thirty days should be able to say what recovery those thirty days buy that five days does not. If the real requirement is a permanent record rather than a recovery buffer, the answer is not a longer bound on an operational store at all — it is a copy of those records in a system built for keeping them. ## What varies between platforms - Where the size bound is per partition of a stream, the arithmetic above divides by the partition count; where it is per stream, it does not; where it is per subscriber backlog, it multiplies by the number of subscribers instead. - Some platforms offer no size bound, so the age bound is the only lever and capacity planning falls entirely to monitoring bytes held. - Where retention is per subscriber rather than shared, "how long may a reader be down" becomes a per-subscriber question and the answer can legitimately differ between them. - Rented clusters often express the same policy in their own units; the derivation is unchanged, only the number you type.
- A team asks for ninety days of retention on an operational stream. How do you respond?Ask what recovery those ninety days buy that the default does not. If the answer is a permanent or auditable record rather than a recovery buffer, the right home is a copy in a system built for keeping records, not a longer bound on an operational store. If it genuinely is recovery, price the capacity and name an owner.
- Why use the peak write rate rather than the average when sizing the byte ceiling?Because the surge is precisely when the history matters and precisely when a ceiling sized on the average starts binding. Sizing on the average guarantees that the retention promise lapses under load, silently, at the moment a reader most likely needs to replay.
- What would make you revisit a stream's bounds without anyone asking?A large growth in its write rate since the bounds were set, a rise in its partition count, or an observed readable span drifting down towards the committed recovery span. All three change what the same configuration means, and none of them looks like a change to anyone reading the settings.
saying these in an interview costs you the question
- Picks a round number of days with no recovery requirement behind it
- Sizes the byte ceiling from the average write rate
- Ignores weekend detection time when estimating a tolerable outage
- Treats the byte ceiling as the retention policy rather than a guard
- Grants a long retention window with no owner and no cost signal
- Uses an operational stream as the permanent record of its events