skip to content

Across many streaming services each team picks its own overflow policy by habit - what would you standardise, and what stays a team decision?

level: principalimportance: should knowfreq 34%

answer

  1. bind the choosing, not the choice
  2. classify the item first
  3. policy recorded next to the code
  4. allow loss only where it is counted
  5. review the exceptions, not the defaults

basics

~20 s

Standardise the decision procedure, not the policy: every stream that may discard must classify what one item means, record the policy it chose and why, and publish a count of what it discards. The policy itself and the buffer size stay with the team.

solid answer

~50 s

A fleet-wide rule that names one policy is wrong on arrival, because the right policy depends on what an item means and that differs stream by stream - the same rule would either discard instructions or buffer snapshots nobody wanted. What generalises is the **procedure**: require every stage that can discard to state the classification of its items (superseded, interchangeable, irreplaceable), to record the chosen policy next to that classification, and to emit a loss count. Add a default for the common case so the easy path is a good one, plus a recorded exception rather than a silent override. Then review the exceptions and the streams whose loss rate never returns to zero. Capacity, thresholds and the policy itself remain local, because only the team knows the burst shape and what a lost item costs.

go deeper

for a junior

The takeaway is smaller than the question: when a stage can throw values away, somebody should have decided that on purpose and written down why. Inheriting the setting from a copied pipeline is how data goes missing.

for a middle

Be able to argue why one fleet-wide policy cannot be right - the policy follows from what an item means, and that differs per stream - and what a standard can require instead without dictating the answer.

for a senior

Bring the operational half: the loss counter as the precondition for permitting discard, the escalation from sustained loss to failure, and what you would alert on across many services rather than one.

for a principal

This is your call to own. Split it as visibility versus decision: the standard owns classification, recording and measurement so the fleet is reviewable; teams own policy and capacity because the knowledge is local. Then measure the standard itself by coverage and by whether the exception register is actually read.

## What a fleet-wide rule can and cannot fix The problem worth solving is not that teams choose differently - they should. It is that most of them are not choosing at all: a stage inherits whatever policy the pipeline it was copied from used, nobody wrote down why, and the loss it produces is unmeasured. A standard that names one policy for everyone replaces an unexamined default with a mandated one and gets the expensive streams wrong. A standard that binds the *choosing* leaves the answers local while making them deliberate. ## Mandate the procedure Three requirements carry almost all of the value, and each is cheap enough to survive contact with delivery pressure: 1. **Classify the item.** Every stage that may discard states what one item represents: a snapshot superseded by its successor, one interchangeable sample among many, or an irreplaceable effect. This is one line, it is the input to every other decision, and a team that cannot answer it has found a real problem rather than a paperwork one. 2. **Record the policy beside the classification.** Not in a wiki page - next to the code, so the reviewer of the next change sees both. The pairing is what makes a wrong policy reviewable: an irreplaceable classification next to a discarding policy is a defect visible without running anything. 3. **Emit a loss count.** A stage permitted to discard must publish how much it discarded, as a rate. Loss you cannot measure cannot be reviewed, alerted on, or argued about with evidence, and a permission to discard without measurement is a permission to lose data indefinitely and unknowably. Notice what these have in common: none of them says which policy is right, and all of them make a wrong policy visible. ## A default with a recorded exception Standards work when the easy path is the good one. Pick the default that fits the majority case in your estimate - in most fleets that is a bounded buffer sized for the expected burst, with a named policy for when it fills - and make departing from it a recorded exception rather than a silent override. The recorded part matters more than the default: an exception register is the list you review, and it concentrates attention on exactly the streams where someone deliberately decided that data may be lost. Resist two temptations. One is mandating fail-fast everywhere as the safe choice: it converts every freshness-driven feed into an outage, and teams will quietly route around a rule that pages them for the map redraw falling behind. The other is mandating a single buffer capacity: burst shapes differ by orders of magnitude between streams, so one number is either wasteful memory or a policy that fires constantly. ## What stays with the team | decision | owner | why | |---|---|---| | the classification of the item | team, reviewed | only they know what the data means | | the overflow policy | team | follows from the classification, which is local | | buffer capacity and thresholds | team | burst shape and memory budget are local | | that a loss count must exist | standard | comparability and reviewability are fleet-wide | | what the counter is named and where it lands | standard | so one query answers the fleet-wide question | | the escalation from sustained loss to failure | standard sets a floor, team tightens | the floor prevents indefinite silent filtering | The split has a shape: the standard owns what makes decisions visible and comparable, the team owns the decision. ## Knowing whether it worked A standard nobody can measure is indistinguishable from a memo. Two signals show whether this one is live. The first is coverage: what fraction of discarding stages publish a loss count, trending toward all of them. The second is the exception register's content - if it is empty, the default is being applied where it does not fit and the rule is being complied with rather than used; if it grows without review, the register is a filing cabinet. The thing you actually want to be able to do, a year in, is answer a question that is unanswerable today: which streams in this organisation are currently losing data, and who decided that was acceptable. ## Both failure modes are real No standard at all leaves loss unmeasured and untraceable, and the cost lands on whoever eventually reconciles a number that does not add up. A standard that dictates the policy produces either outages on feeds that should have shed data or silent loss on feeds that should not, plus a culture of working around the rule. Between them sits the narrow thing worth writing: make the choice explicit, make its consequence measurable, and leave the choice itself where the knowledge is.

  • A team argues that emitting a loss count for every discarding stage is overhead they should not pay. How do you answer?
    The cost is a counter increment on a path that is already shedding work, which is the cheapest moment in the system to spend one. The alternative is a stage whose output is a filtered version of its input with no record of the filter. If the loss genuinely never matters, the counter stays at zero and costs nothing; if it is ever non-zero, it is the only evidence that will exist.
  • What would make you tighten the standard from a recorded exception to an approval?
    Evidence that exceptions are being recorded but not read - a register that grows while the loss counters on those same streams sit non-zero for weeks. At that point the lightweight rule has stopped functioning as a review mechanism, and the material consequence, data lost from streams somebody flagged as unusual, is worth a second pair of eyes before it ships rather than after.

saying these in an interview costs you the question

  • Mandates one overflow policy for every stream in the organisation
  • Leaves the choice entirely local with nothing recorded anywhere
  • Believes a standard can work without loss being measured
  • Fixes one buffer capacity for streams with different burst shapes
  • Treats an empty exception register as proof the standard is working