skip to content

A writer discards records instead of failing the send when its unsent buffer fills under cluster saturation — what has that traded?

level: seniorimportance: should knowfreq 44%

answer

  1. three exits, all with a price
  2. discard keeps the caller healthy
  3. nothing fails, so nothing signals
  4. statistical value against business fact
  5. reconcile offered against accepted

basics

~20 s

Availability of the calling path has been bought with data, and with the evidence that the data is gone. Sends keep succeeding, error rates stay flat, and the loss is detectable only by reconciling records offered against records the cluster accepted.

solid answer

~50 s

Of the three things a client can do with a full **writer send buffer** — block the caller, fail the call, discard the record — discarding is the only one that keeps the application entirely healthy. That is the trade: the producing service never stalls and never errors, and in exchange records vanish with no signal at the call site. It is defensible where a record's value is statistical rather than individual and where a gap in the series is tolerable — sampled measurements, progress pings, chatter emitted far faster than anyone reads it. It is indefensible wherever a record is the only evidence that something happened, because the application will have told a user the operation succeeded. The rule of thumb: discard only where you would also accept a gap in the series being invisible for a week.

go deeper

for a junior

Recall that a client with nowhere to put a record can drop it and still report success, so a send that returns without error does not prove the record reached the cluster.

for a middle

Explain the three exits from a full writer buffer and what each costs, and be able to say which kinds of record tolerate a gap and which do not.

for a senior

Show that you would set this per stream rather than per service, count the discards, and reconcile records offered against records accepted before trusting that nothing was lost.

for a principal

Treat it as a policy question: which classes of record in the estate may ever be discarded, who is allowed to decide that, and what evidence exists afterwards that the decision was honoured.

## What the discard option actually is When a node is saturated it stops draining the writer as fast as the writer fills, and the client's **writer send buffer** reaches its ceiling. At that instant the client has exactly three exits, and they are mutually exclusive for any given record: - **Block** the calling thread until space appears — availability of the record, paid for with availability of the application. - **Fail** the send — availability of the application, paid for with an error the application must now handle. - **Discard** the record and let the call succeed — availability of both, paid for with the record and with any knowledge that it existed. The third is known as *the quiet drop* because it is the only one with no symptom. Nothing throws, no counter of failures moves, no thread is parked, and the operation the user asked for completes normally. ## What is traded | Exit | Application stays up | Record survives | Loss is visible | |---|---|---|---| | Block the caller | No | Yes, while the process lives | Yes — the stall is obvious | | Fail the send | Yes | Only if the code re-routes it | Yes — an error per record | | Discard the record | Yes | No | No, not at the call site | The row that matters is the third column against the fourth. A failure is an invitation to do something: retry later, hold the record somewhere durable of your own, fail the user's operation honestly, or shed deliberately and count what you shed. A discard removes the invitation, and with it the option. The decision about whether this record mattered has been made in advance, by configuration, identically for every record the writer ever sends. There is a second-order cost that is easy to miss. Because nothing fails, nothing signals the saturation upstream either. A writer that fails would slow down, back off, or complain; a writer that discards keeps offering exactly the same load to an already-saturated node, so the node stays saturated and the loss continues for as long as the burst does. ## When discarding is defensible It is a legitimate choice, not a mistake, in a recognisable shape of workload: - The record's value is **statistical** — one sample among many of the same thing, where a conclusion drawn from 98% of the series is the same conclusion. - The record is **superseded quickly** — a periodic position report, a heartbeat, a current-value update where the next one is along shortly and carries the same truth. - The volume is **far above what anyone consumes**, so the system was already sampling in effect. - Losing the record **costs less than stalling the producer**, which is the case when the producer is doing something more important than producing that record. In those cases, blocking the caller would be the actual defect. ## When it is not - The record is the **only evidence** that a business event occurred, and no other store holds it. Anything the application has already told a user it did falls here. - The record **drives an action elsewhere** — a payment, a shipment, a notification — so the loss becomes someone else's missing outcome rather than a gap in a chart. - The record is **required to be kept**, for audit or regulatory reasons, in which case a configuration setting has quietly become a compliance decision. - The stream is **reconciled against another system**, where a silent gap turns into a mismatch investigated days later at far greater cost than an error would have been. ## Making the choice explicit Three habits make this an engineering decision rather than an inherited default: 1. **Decide it per stream, not per application.** One service frequently emits both measurements and business facts; if it has a single writer, the weakest record's policy is applied to the strongest one. Two writers with different policies is a small price. 2. **Count what you discard.** A client that drops records almost always exposes a count of them. That count is the only cheap evidence the loss exists, and it belongs somewhere a human sees. 3. **Reconcile.** Periodically compare records the application handed to the client against records the cluster accepted for the same stream and interval. A persistent difference is either discards or an unsent buffer that never drained; both are worth knowing about. The interview answer worth giving is not "discarding is bad". It is that the three exits are a choice between the application, the data and the evidence, that discarding gives up two of the three, and that the choice should be made where the value of the record is known — not by whichever default the client shipped with.

  • How would an operator notice the discard at all?
    Two cheap sources. The client usually keeps a count of records it dropped, which can be surfaced like any other application number. Failing that, reconcile: compare how many records the application handed to the client against how many the cluster accepted for the same stream over the same interval, and investigate any standing difference.
  • Which records are never safe to discard?
    Any record that is the only evidence an event happened, anything the application has already confirmed to a user, anything that triggers an action elsewhere such as a payment or a shipment, and anything retained for audit. In those cases the honest exits are failing the send or holding the record durably yourself.
  • Why does a discarding writer prolong the saturation it is reacting to?
    Because nothing about it slows down. A writer that fails or blocks reduces the load it offers, which gives the node a chance to recover; a writer that discards keeps offering the identical rate, so the node stays over capacity and the discards continue for the whole duration of the burst.

saying these in an interview costs you the question

  • Calls discarding acceptable because no error was reported
  • Assumes blocking the caller is always the safer default
  • Applies one buffer-full policy to every stream a service writes
  • Forgets that a discarding writer keeps loading an already saturated node
  • Expects cluster-side dashboards to reveal records that never arrived