skip to content

An age bound of twenty-four hours is configured, yet records forty hours old are still readable — why is nothing broken?

level: middleimportance: should knowfreq 50%

answer

  1. eligible, not deleted
  2. whole segments, never single records
  3. the active one is exempt
  4. youngest record keeps a segment alive
  5. a sawtooth, not a line

basics

~20 s

Enforcement is coarse. Stores that append into rolled segments remove a whole closed segment at a time, never a single record, and the segment still being written is never removed — so surviving history always runs somewhat past the bound.

solid answer

~50 s

An age bound makes records **eligible** for removal; it is not a deadline by which they disappear. On a store that appends into segments, the unit of removal is a whole closed segment, so a segment survives until everything in it qualifies, and the **active segment** — the one still being appended to — is never removed at all. A quiet stream that rolls a segment once a day can therefore hold far more than its bound. The removal pass also runs periodically rather than continuously, adding its own delay. The practical consequences: the surviving history is always a little more than configured, the earliest position still on the store moves up in **jumps** rather than smoothly, and you must never promise anyone that a record is gone at a particular hour on the strength of an age bound.

go deeper

for a junior

Know that a retention age says records may be removed after that long, not that they vanish on the hour. Seeing slightly older records is normal, not a bug.

for a middle

Explain the segment as the unit of removal, why a segment survives on its youngest record, and why the active segment is exempt — that trio is the whole answer.

for a senior

Bring the operational consequences: the readable span is a sawtooth whose low point is what you can rely on, and a quiet stream is the case that surprises teams in production.

for a principal

Make the policy point: an age bound is a capacity and replay instrument, never a mechanism for promising anyone that data has ceased to exist by a date.

## An age bound is an eligibility rule The most useful correction a candidate can make here is to stop calling the age bound a deadline. What it actually says is: *a record older than this may be removed.* It does not say when removal happens, and it certainly does not say a record becomes unreadable at that moment. Between eligibility and disappearance sit three mechanisms, all of them coarse. ## 1. Removal works on whole segments Stores in this class append records into **segments** — files that are written until they reach a size or an age, then **rolled closed** and never appended to again. Removal operates on that unit: a closed segment is dropped in one piece. That has an immediate consequence. A segment is only removable once everything in it is past the bound — which means it survives on the strength of its **youngest** record. A segment holding records from 10:00 to 14:00, with a twenty-four-hour bound, is not removable until 14:00 the next day, at which point its 10:00 records are twenty-eight hours old and have been readable the whole time. Removing individual records would require rewriting the segment, which is exactly what an append-only store is designed not to do. ## 2. The active segment is never removed The segment currently being appended to is not a candidate. On a busy stream this hardly matters, because segments roll frequently. On a **quiet** stream it matters enormously: if the roll is triggered by size and the traffic is a trickle, the active segment may stay open for days, and every record in it stays readable regardless of the age bound. This is the usual explanation for a low-traffic stream that appears to ignore its retention policy entirely. Most stores also allow a roll to be forced by age for exactly this reason, but the effect remains: the bound binds on closed segments only. ## 3. The removal pass is periodic The background **removal pass** that enforces the bounds runs on an interval. Between runs, eligible segments are still on the data volume and still readable. The interval is small compared with a typical age bound, but it is not zero, and it stacks on top of the two effects above. ## What this does to the front of the history Put together: the **earliest position still on the store** — the floor under everybody, below which nothing can be read — does not creep forward continuously. It sits still while a segment fills, then jumps forward by a whole segment when that segment is dropped. A chart of readable span against time is a sawtooth, not a line. Anyone reasoning about the retention window (the span of history still readable, which is the replay budget a recovering reader spends) has to reason about it as a sawtooth whose **low point** is what they can rely on, not its average. So the honest statement of the guarantee is one-sided: - **You may rely on:** history reaching back *at least* the age bound, minus the small slack of a pass interval — and in practice rather more. - **You may not rely on:** a record being gone once it passes the bound. ## Which clock the age is read from One more coarse edge, and the one that surprises people most. The age is compared against **the record's clock**, and stores differ on which clock that is: the timestamp the writer supplied with the record, or the time the record arrived at the store. Where the writer's timestamp is used, a backfill of genuinely old records becomes removable the moment it lands — a replay of last year's data, written into a stream with a thirty-day age bound, can vanish within one removal pass, having never been readable long enough for anyone to consume it. The same mechanism also means a writer with a badly skewed clock can extend or destroy retention for everything in the segments it writes into. ## What varies between platforms - Segment-based removal is the norm for stores that keep a shared append-only history; designs that track a per-subscriber backlog can and do remove records individually as each subscriber finishes with them. - Whether the age is measured from a writer-supplied timestamp or from arrival differs by platform and is sometimes configurable per stream. - Roll triggers differ — size, elapsed time, or both — which is why the same bound behaves differently on a busy and a quiet stream. ## What interviewers listen for - "Eligible for removal", not "deleted at". - The whole-segment unit, and the segment surviving on its youngest record. - That the active segment is exempt, and why a quiet stream is the pathological case. - The awareness that an age bound is not a compliance mechanism for making data go away.

  • Which clock is the age measured from, and why does it matter for a backfill?
    Either the timestamp the writer supplied or the time of arrival at the store, depending on the platform and sometimes on the stream's own setting. If the writer's timestamp is used, a backfill of year-old records is instantly past a thirty-day bound and can be removed by the next pass, effectively deleting itself on arrival.
  • Can an age bound be used to promise that records are unreadable after a fixed period?
    No. It makes records eligible for removal, and coarse enforcement means the real disappearance is later and imprecise — sometimes much later on a quiet stream. A hard 'gone by this date' requirement is a different problem with different machinery, not a retention setting.
  • Why does a nearly idle stream appear to ignore its retention policy?
    Its active segment rarely rolls closed, and only closed segments are removable, so records stay readable well past the bound. Forcing a roll on elapsed time rather than size alone brings behaviour back in line, at the cost of more, smaller segments.

Rubbish is collected by the skip, not by the item. A skip is only hauled away once it is closed and everything inside is old enough — and the one still being filled is never hauled at all, however old the thing you dropped in first.

saying these in an interview costs you the question

  • Treats the age bound as a deadline by which records are gone
  • Thinks records are removed one at a time as each passes the bound
  • Forgets that the segment being appended to is never removed
  • Assumes a segment is removable as soon as its oldest record qualifies
  • Believes the age is always measured from arrival at the store
  • Offers an age bound as a way to guarantee data is unreadable by a date