A stream quietly discards values whenever load spikes and every dashboard still looks healthy - how do you make that loss visible?
answer
- absence leaves no trace downstream
- depth and throughput both look healthy
- count at the point of discard
- attach a gap marker to survivors
- sustained loss should become a failure
basics
~20 sCount every discard at the point it happens, publish it as a rate beside throughput, and carry a gap marker with the next surviving value so consumers can see what they missed. Escalate sustained loss into a failure rather than leaving it silent.
solid answer
~50 sSilent discard is invisible by construction: the discarded value leaves no trace anywhere downstream, and the two obvious signals actively conceal it - a discarding stage holds buffer depth pinned at its ceiling and output rate pinned at the consumer's capacity, which is exactly what a healthy stage looks like. So instrument the only place where the information still exists, the discard itself: increment a counter there and expose it as a **rate**, not a lifetime total, so a burst is visible against a background of zero. Attach a gap count to the next value that is admitted, so a consumer can tell that it is looking at a sampled view rather than a complete one. Then alert on the loss rate and decide a threshold beyond which the stage should stop discarding and fail instead.
code
pseudocode · 8 linesfunction onArrival(value):
if count(buffer) == capacity:
droppedSinceLastAdmit = droppedSinceLastAdmit + 1
incrementCounter(stageName, "values.dropped")
return
append(buffer, { payload: value, skippedBefore: droppedSinceLastAdmit })
droppedSinceLastAdmit = 0go deeper
Understand the core trap: a value that was dropped leaves nothing behind, so nobody downstream can notice it. If a stage may discard, something has to count the discards where they happen.
Explain why depth and throughput both look healthy during discarding, and describe the counter plus gap-marker technique well enough to implement it, including why a rate beats a lifetime total.
Bring the operational shape: what you alert on, how you tell a burst that recovered from a permanent mismatch being hidden, and the threshold at which you would rather the stage failed than kept filtering reality.
Make instrumented loss a precondition for permitting a discarding policy at all, and treat a stage that has discarded continuously for weeks as an undeclared change to the product's data, not as a tuning issue.
## Silent loss is invisible by construction A value that is discarded at a full buffer is never delivered, never logged by the consumer, and never missed by anything that did not know it was coming. There is no downstream observer that could reconstruct it, which is what separates this failure from almost every other production problem: the usual move of looking harder at the consumer finds nothing, because from the consumer's side a lossy stream and a quiet producer are identical. Worse, the two metrics a team reaches for first are precisely the ones the policy flattens. | signal | what it shows | what it hides | |---|---|---| | buffer depth | how full the stage is right now | discarding pins depth at capacity, so a steady line means either idle or saturated | | output rate | values the consumer processed | it equals consumer capacity under overload, which is what a healthy busy stage also shows | | consumer latency | how long the survivors waited | says nothing about arrivals that were refused before they ever waited | | error rate | failures that were signalled | a discarding policy signals nothing, so this stays flat by design | That table is the whole diagnosis: none of the default signals can distinguish a stage that lost nothing from one that lost most of its input. ## Instrument the discard itself The information exists at exactly one instant - the moment the policy refuses a value - and nowhere afterwards. Three things are worth recording there: - **A counter of discarded values**, incremented on every drop. This is the primary signal and it costs an increment. - **The rate**, derived from the counter, published beside throughput. A lifetime total is nearly useless: it grows monotonically and nobody can tell yesterday's incident from this minute's. A rate that sits at zero and spikes is legible at a glance. - **Enough identity to aggregate**, usually the stage name and, where a stream is keyed, the key. A fleet feed that drops everything for one vehicle is a different defect from one that drops a little for all of them, and an undifferentiated counter cannot tell them apart. What is not worth doing is logging each discard as an event. The condition fires precisely when the system is at its busiest, so a per-value log line turns a data-loss problem into a log-volume problem and often makes the overload worse. ## Carry the gap to the consumer A counter tells the operator. It does not tell the consumer, which may be computing something whose correctness depends on completeness. The cheap technique is to carry the gap with the data: keep a running count of values discarded since the last admitted one, attach it to the next value that is admitted, and reset it. The consumer then receives a sequence in which every element states how much was lost immediately before it, and can decide for itself - suppress an aggregate, mark a display as degraded, or reconcile later. This is the difference between a stream that is lossy and a stream that is *honestly* lossy. ## When counted loss should become a failure Counting makes loss visible; it does not make it acceptable. The policy decision that follows is a threshold: how much loss, sustained for how long, means the stage should stop discarding and fail instead. Two anchors help: 1. **Bursty loss with a returning-to-zero rate** is the policy working. The buffer filled, some values went, the backlog drained. Alert on the pattern only if it is new or growing. 2. **Sustained loss with a rate that does not return to zero** is not overload absorption, it is a permanent rate mismatch being hidden. A stage in that state is producing a filtered version of reality indefinitely, and failing loudly is usually more honest than continuing, because a failure reaches someone who can act while a steady discard rate on a chart that nobody opens reaches nobody. ## Closing the loop The review question that catches this class before production is short: for every stage that may discard, where is the counter, and who is paged when it is non-zero for longer than the burst it was sized for? A discarding policy without that answer is not a policy, it is an undocumented data filter, and its cost is paid by whoever eventually reconciles the numbers.
- Why is a lifetime total of discarded values a poor signal compared with a rate?Because it only ever grows, so it cannot distinguish an incident last month from one happening now, and any alert on it fires forever once tripped. A rate sits at zero when the stage is healthy, spikes visibly during a burst and returns, which makes both the event and its recovery legible - and makes a rate that fails to return the distinct, alertable condition.
- The consumer computes an average from a feed that sometimes discards. What should it do with the gap count?Treat it as a confidence signal rather than ignoring it. Loss under overload is not random - it happens during spikes, so the surviving sample is biased toward calm periods and the average flatters reality. The honest handling is to publish the average alongside the loss it was computed over, or suppress it entirely while the gap count is non-zero.
saying these in an interview costs you the question
- Reads a flat buffer-depth chart as proof that nothing was lost
- Counts discards but never alerts on the rate
- Assumes downstream consumers will notice missing values by themselves
- Logs every discarded value individually under overload
- Treats an indefinitely lossy stage as acceptable because the service stays up