A vehicle-position feed and a settlement-instruction feed both outrun their consumers - why can one discard values and the other not?
answer
- ask what one item means first
- snapshot, sample, or irreplaceable effect
- superseded values cost nothing to lose
- an effect cannot be re-derived downstream
- one service, several policies
basics
~20 sClassify the item first. A position is a snapshot that the next one supersedes, so discarding intermediates costs the consumer nothing. A settlement instruction has an effect of its own that no later item re-derives, so discarding it is a lost transfer.
solid answer
~40 sThe policy follows from what a single item means, not from a default. A vehicle position is a **snapshot of a state**: the next one replaces it, so keeping only the latest and throwing away the intermediates delivers exactly what the map needs, and a deep buffer would only draw the vehicle where it used to be. A settlement instruction is an **event with an effect**: nothing later in the stream reconstructs the transfer a dropped instruction would have caused, so the discarding policies are simply wrong there - the honest options are buffering while the backlog drains and failing loudly when it does not, so that someone is told rather than silently short-changed. The interview move is to ask what one item represents before naming any policy.
go deeper
Remember the two shapes of data: items that replace each other, like a current position, and items that each do something, like an instruction. Only the first kind can be thrown away without losing anything.
Be able to classify a feed as superseded, interchangeable or irreplaceable and derive the policy from that classification, including why a deep buffer on a freshness-driven feed produces complete data and a wrong display.
Show the failure you would expect from the wrong pairing and how it surfaces: a map drawing stale routes after a burst, or a reconciliation gap weeks later from a discarding stage copied onto an instruction feed.
Own the decision procedure rather than the answer. Make item classification a required step when a stage declares an overflow policy, and accept that one service will legitimately run several different policies at once.
## The question behind the question Asked which overflow policy to use, a weak answer names one. A strong answer asks a question back: what does a single item in this stream represent? The four policies are not ranked in the abstract - each is correct for some data and indefensible for other data, and the classification of the item is what decides. Two feeds running side by side in the same service make the point sharply. ## Three kinds of item 1. **A snapshot of a state.** Each item states the current value of something that keeps changing - a position, a price, a gauge reading, a progress percentage. Item *n+1* supersedes item *n* completely. Discarding intermediates loses no information the consumer can use, because the consumer only ever wanted the current value. 2. **A sample from a population.** Each item is one observation among many and the consumer aggregates them - a rate, a percentile, a rolling average. Individual items are interchangeable; losing some biases the aggregate, which may or may not matter, but no single loss is a correctness defect on its own. 3. **An event with an effect.** Each item causes something to happen exactly once - an instruction, a command, a record that must be persisted. Nothing later re-derives it. Discarding one is not a degraded view; it is a wrong outcome. Only the third class forbids discarding outright. The first class positively *wants* it. ## The position feed The fleet map redraws vehicles from a feed that emits far faster than the renderer can draw. Keeping only the latest position per vehicle is not a compromise here - it is the correct behaviour. Consider the alternative: a generous buffer preserves every position, and after a burst the map faithfully draws each vehicle along the route it travelled several minutes ago. The data is complete and the display is wrong. Dropping the newest arrival is worse still, because it keeps precisely the stalest positions and refuses the current one. The value of an item in this feed decays to nothing the moment its successor exists, so freshness is the thing the policy must protect and superseded values are what it may spend. ## The instruction feed The settlement feed carries instructions that move money when processed. Here every discarding policy fails for the same reason: the effect of an item exists nowhere else. Keeping only the latest instruction means the transfers it superseded never happen, and nothing downstream is able to notice, because the consumer sees a shorter but perfectly well-formed sequence. The defensible options are to buffer while the backlog drains - accepting memory and latency as the price of losing nothing - and to fail loudly when the backlog will not drain, so that a human or a supervising component learns that the pipeline could not keep up. Failing costs availability, which is expensive; silent loss costs correctness, which is worse and much harder to discover. ## One pipeline may hold both The policy is a property of the stage and the data, not of the service, so a single application often runs several: | feed | item means | policy | what it spends | |---|---|---|---| | vehicle positions | a snapshot superseded by the next | keep only the latest | superseded positions, which nobody wanted | | settlement instructions | an effect that happens once | bounded buffer, then fail | memory and latency, then availability | | telemetry samples for a rolling average | one observation among many | drop under pressure, with the loss counted | statistical fidelity, knowingly | A team that sets one overflow policy for a whole service therefore gets at least one of its streams wrong, and usually the expensive one. ## How this goes wrong in practice The usual path is inheritance. A pipeline is built for the map feed, the discarding policy is right there, and the instruction feed is later added by copying the working stage. Nothing fails, because silent discard looks exactly like healthy operation, and the defect surfaces weeks later as a reconciliation difference nobody can trace to a stream. The defence is a habit, not a tool: whenever a stage declares an overflow policy, the reviewer asks what one item on that stage represents, and expects an answer in the vocabulary above - superseded, interchangeable, or irreplaceable. One caveat worth saying aloud: if the items are irreplaceable and the mismatch is sustained rather than bursty, no in-stream policy is a good answer. Buffering only postpones the failure and failing only announces it. The real fix is upstream or sideways - make the consumer faster, partition the work, or move the items out of a live stream altogether - and choosing an overflow policy is what you do while that remains true.
- The instruction feed cannot discard and cannot afford an outage either. What is left?Nothing inside the stream, which is the useful answer. Buffering postpones and failing announces; neither adds throughput. The remaining moves are to raise consumer capacity, partition the feed so several consumers share it, or take the items out of a live pipeline and hand them to something durable. An overflow policy is what protects you until one of those lands.
- Telemetry samples feed a rolling average. Is discarding them under pressure safe?Safe for correctness, not for accuracy. No single sample is irreplaceable, so nothing breaks, but the loss is rarely uniform - it happens exactly during load spikes, so the surviving samples are biased toward calm periods and the average flatters the system. Discard if you must, count what you discarded, and treat the metric as untrustworthy while the count is non-zero.
saying these in an interview costs you the question
- Picks an overflow policy from a default rather than from the data
- Cannot say what a single item in the stream represents
- Believes keeping only the latest is safe for items with side effects
- Assumes a rare drop is acceptable in any stream
- Sets one overflow policy for every stream in a service