When should a ring buffer overwrite its oldest record instead of rejecting the newest write?
answer
- ask what the reader does with it
- which end of the history matters?
- crash diagnosis versus a billing ledger
- silent loss is the actual defect
- count the drops, number the records
basics
~20 sOverwrite when recency beats completeness — a crash-diagnosis recorder wants the last N events whatever came before. Reject when every record must survive, such as billed usage or an audit trail. Either way, count the losses and expose them.
solid answer
~50 sIt is a product decision, not an implementation detail, and it turns on which end of the history the reader needs. A flight-recorder-style buffer holding the last N events of a running service is read *after* something goes wrong, so the newest records are the valuable ones and overwriting the oldest is the feature — it also keeps the recording path allocation-free and incapable of failing. Reject-newest is right when the oldest records are the committed ones: billed usage, audit trails, anything a downstream system reconciles. What is never acceptable is losing silently. Whichever policy you pick, keep a drop counter and put a sequence number on each record so a consumer can detect a gap and size it. And do not log per drop: under the overload that fills the buffer, that amplifies the very problem you are observing.
go deeper
Know that a fixed-capacity buffer must do something on a full write, and that the two choices are discarding the oldest stored record or refusing the new one. Be able to name one use for each.
Explain how each policy is implemented and what it does to the indices, and why a discarded record needs a counter and a sequence number if anyone is to notice it went missing.
Pick the policy from the reader's needs and defend it: recency versus completeness, the never-failing recording path, the lapped-consumer hazard, and sizing from peak rate times the diagnosis window.
Own it as policy across a fleet: a memory ceiling per instance, a defined loss budget, the requirement that any dropping component export its drop rate, and the call about which data streams are simply not allowed to live in a lossy in-memory structure at all.
## The decision is about the reader, not the buffer A fixed-capacity buffer will eventually receive a write with nowhere to put it. There are exactly two in-structure answers — discard the oldest stored record to make room, or refuse the incoming one — and choosing between them requires knowing who reads the buffer and what they do with it. **Overwrite-oldest** suits the flight-recorder pattern: a small in-memory buffer that continuously records the last N timestamped events of a running service, is never read in the normal case, and is dumped when something fails — a crash, a health-check failure, an operator asking what just happened. Here the value of a record decays with age. A record from four hours ago cannot explain a failure ten seconds ago. Overwriting is what keeps the window pinned to the present, and it buys two operational properties that matter more than completeness: the recording path never fails, and its memory footprint is a constant you can multiply by the fleet size and put in a capacity plan. **Reject-newest** suits the opposite shape: records that are individually owed to someone. Billed usage, entries in a ledger a downstream system will reconcile, security-relevant audit events. Overwriting one of those to make room for a fresher one destroys a committed fact, and a fresher fact is not a substitute. Here the correct behaviour on a full buffer is to refuse, report the refusal to the caller, and let a layer that can afford to persist deal with it. If both answers are unacceptable — you can neither lose the old nor refuse the new — then a fixed-size in-memory ring is simply the wrong structure for that data, and the honest answer in an interview is to say so rather than to invent a third policy inside the ring. ## Make the loss observable The genuine defect is not dropping. It is dropping invisibly. Two mechanisms cover it: - **A drop counter.** A monotonically increasing count of records the buffer discarded or refused, exported like any other metric. It converts a silent correctness problem into a visible rate, and its derivative tells you whether you are under-sized or facing a burst. - **Per-record sequence numbers.** If each record carries the sequence value it was written at, any consumer that remembers the last value it saw can compute exactly how many records it missed. That turns "the data looks weird" into "I missed 412 records between these two timestamps", which is the difference between a usable recording and a misleading one. Sequence numbers also solve the **lapped-reader** hazard specific to overwrite-oldest. A consumer holding a cursor into the buffer can fall far enough behind that the writer laps it, and the slot the cursor points at no longer holds the record it expected. Comparing the cursor against the current write sequence detects this before the read: the consumer is told it was lapped and by how much, instead of quietly returning records from the wrong era. ## The amplification trap A tempting instinct is to log a warning on each drop. Do not. Buffers fill under load, and under load the drop rate is high, so a per-drop log line multiplies exactly when the system is least able to absorb it — and if that log line is itself recorded through a buffered path, the recorder is now recording its own distress. Count the drops, report the count on a fixed cadence, and log at most a rate-limited summary. ## Sizing Size a recorder by the window you need to reconstruct, not by a round number: peak event rate multiplied by the window in seconds multiplied by the record size, with headroom, checked against the memory ceiling per instance and then across the fleet. Use the peak rate, not the mean — the burst is when you need the recording. If the arithmetic does not fit the ceiling, shrink the *record* before you shrink the window: drop fields, or admit only records above a severity threshold and sample the rest. A recorder whose window is shorter than the incident it must explain is memory spent for nothing. ## Middle grounds worth naming Severity-aware admission — always keep errors, sample the routine records — keeps the window long for the events that matter. Two rings at different resolutions, one fine and short, one coarse and long, give both a detailed recent view and a longer trend. And a trigger-driven snapshot, where the buffer is dumped on a defined condition, converts a rolling window into an artifact someone can read later. ## What separates a strong answer The weak answer treats every dropped record as a bug to be fixed. The strong one names the reader, names the window, picks the policy that follows from those two, and then explains how anyone downstream will know a loss happened at all.
- How does a consumer of an overwrite-oldest recorder learn that it missed records?Give every record the sequence number it was written at and let the consumer remember the last one it read. Before each read it compares its cursor to the writer's current sequence: if the difference exceeds capacity it was lapped, and the gap size is the exact number of records lost. Without sequence numbers the consumer cannot distinguish fresh records from ones it has already seen, and reports a plausible but wrong history.
- How do you size a fixed-capacity event recorder?Peak event rate times the window you must be able to reconstruct times the record size, plus headroom — then check it against the per-instance memory ceiling and multiply by the fleet. Use peak rate rather than mean, because bursts are exactly when the recording matters. If it does not fit, shrink the record or admit only higher-severity events rather than quietly shortening the window below the length of the incidents you need to explain.
- Why is emitting a warning per rejected record often worse than either policy?Because the warning is itself an event, and the buffer only rejects under the load that produced the flood, so the per-record warning multiplies precisely when the system is most stressed — and if warnings travel through a buffered path, the recorder starts recording its own overflow. Keep a drop counter, publish it on a fixed cadence, and rate-limit any human-readable summary to one line per interval.
saying these in an interview costs you the question
- Treats every dropped record as automatically a defect
- Overwrites audit or billing records to make room
- Drops silently with no counter and no sequence numbers
- Logs a line per dropped record under overload
- Sizes the buffer by a round number rather than rate times window
- Ignores that a slow consumer can be lapped by the writer