skip to content

A three-day traffic surge left a stream holding nineteen hours of history instead of its usual seven days, with no error raised — what happened and what should have been watched?

level: seniorimportance: must knowfreq 62%

answer

  1. the configuration did not change
  2. bytes fixed, rate up, span down
  3. removal is success, not error
  4. watch the observed span
  5. raising it back returns nothing

basics

~20 s

The byte ceiling began binding before the age bound: at several times the normal write rate the same bytes buy far less time. Removal was correct and silent, so the only signal is the readable span itself, which nobody was measuring.

solid answer

~40 s

Nothing failed. The stream has both bounds live, and at normal rates the seven-day age bound bound first; the surge raised the write rate until the byte ceiling filled in nineteen hours, and from then on **whichever binds first** was the ceiling. Removal of the oldest records is the store doing its job, so there is no error, no alert and no trace beyond the history that is no longer there. The signal to watch is not the configured number but the derived one: **the age of the earliest position still on the store**, which is the retention window you actually have, alerted whenever it falls below the recovery span you promised readers. Raising the ceiling afterwards protects the future and returns nothing; whatever was removed is gone from this store.

go deeper

for a junior

Take away one idea: the number of days in the configuration is not the history you have when traffic rises, because a byte ceiling is also in force.

for a middle

Explain the arithmetic that caused it — a fixed ceiling divided by a much higher write rate — and why the store raises no error for doing exactly what it was configured to do.

for a senior

Show the operating response: the derived span as the monitored signal, an ordered set of actions with capacity checked before the ceiling is raised, and telling the affected readers.

for a principal

Set the rule for the estate: size the ceiling so it never binds in normal operation, and treat the day it starts binding as the day the retention promise stopped being real.

## What actually happened The stream was configured with a seven-day age bound and a byte ceiling. At the ordinary write rate the ceiling implied a span comfortably longer than seven days, so the age bound was the one doing the work and the ceiling sat behind it as a guard on the data volume. The surge changed one variable: bytes per hour. The ceiling is a fixed number of bytes, so the span it implies is `ceiling ÷ current byte rate` — and multiplying the rate divides the span. At roughly nine times the normal rate, seven days of allowance becomes about nineteen hours. From that moment the ceiling was binding, the removal pass was removing the oldest records to stay under it, and the readable history was being consumed from the front as fast as it was being written at the back. The key sentence for an interview: **the retention window is a consequence of throughput, not a promise of the configuration.** ## Why nothing alerted There are three reasons this is invisible, and a good answer names them: 1. **Removal is success, not failure.** The store removed records because it was told to. There is no error condition to report. 2. **Every configured number still reads correctly.** Anyone inspecting the stream sees "seven days" and believes it. The configuration did not change; its meaning did. 3. **The victim is absent.** The loss lands on whoever tries to read back more than nineteen hours — and they usually try days later, during an unrelated incident, which is the worst possible moment to discover the replay budget was spent. ## The signal that would have caught it Monitor the **derived** quantity, not the setting. Two forms, both useful: - **Observed span** — the age of the earliest position still on the store, per partition of a stream. This is the retention window you actually have right now, and the retention window *is* the replay budget. Alert when it falls below the span you committed to readers, with a warning threshold well above it. - **Projected span** — bytes held in the scope divided by the current byte rate. This turns while the surge is still ramping, hours before the observed span has fallen, which is the difference between a capacity decision and a post-mortem. Both are cheap. Neither is on a default dashboard, because a default dashboard shows what the store is doing, not what the store can still give you. ## What to do in the moment In order, and the ordering is the answer: 1. **Decide what the span is worth today.** If a reader is currently behind, or an investigation is in flight, the surviving history is urgent; if nothing needs it, this is a capacity item and not an incident. 2. **Raise the byte ceiling if the data volume genuinely has room**, so the age bound binds again. Confirm the headroom first — a ceiling raised past what the volume can hold converts a retention problem into a much worse one, and that failure is not one you want to meet by accident. 3. **Reduce what is arriving**, if the surge is a misbehaving writer rather than real demand — duplicate traffic and a retry storm are common causes. 4. **Accept and record the loss** of the span already removed, and tell the teams that read this stream, because their assumptions about how far back they can go are now wrong. What does *not* work is raising the ceiling to recover history. Removal on this store is destructive: raising a bound changes what is kept from now on and returns nothing. If the same records exist somewhere else, recovering them is that system's job. ## The lesson to state out loud A byte ceiling should be sized so that in normal operation it **never binds** — it is a guard, and the age bound is the policy. The day the ceiling starts binding is the day your stated retention policy silently became fiction, and the only thing that tells you is a number nobody configured. ## What varies between platforms - Where the size bound is per partition of a stream, a surge that is skewed towards a few partitions truncates only those, so the readable span differs across one stream. - Where retention is a per-subscriber backlog, the same surge shortens the span for the subscribers that are behind rather than for the stream as a whole. - Some platforms expose the age of the earliest position still on the store directly; where they do not, it can be derived by reading the oldest available record's timestamp. - On platforms offering no size bound at all, this failure mode does not exist — the surge grows bytes instead, and the risk moves to the data volume.

  • The ceiling is raised the next morning. What comes back?
    Nothing. Removal is destructive on this store, so a higher ceiling only changes what is retained from that point on. The span that was removed during the surge is unrecoverable here; if those records exist in another system, recovering them is that system's problem, not a retention setting.
  • What single number would you put on the dashboard for this stream?
    The age of the earliest position still on the store, per partition of a stream — the span of history actually readable. Alert when it drops below the recovery span committed to readers. Add bytes held divided by the current write rate as the leading indicator, since it moves before the span does.
  • Why is it a design smell if the byte ceiling binds in normal operation?
    Because the ceiling is a guard on capacity, not a retention policy. If it binds routinely, the stated age bound is fiction: the real retention is whatever today's traffic allows, it changes without anyone editing it, and no team reading the stream can plan a recovery against it.

saying these in an interview costs you the question

  • Treats the shortened history as a malfunction of the removal pass
  • Expects an alert, since the store did something destructive
  • Thinks raising the ceiling afterwards restores removed records
  • Monitors the configured bounds instead of the readable span
  • Assumes the age bound protects history whatever the traffic
  • Raises the ceiling without first checking the data volume has room