skip to content

When a Go ingest pipeline's bounded channel stays full, how do you decide between stalling upstream and dropping events?

level: principalimportance: should knowfreq 30%

answer

  1. who owns the data, not the code
  2. what does the source do when you stop reading?
  3. one policy per class of event
  4. a bigger buffer is not a third option
  5. agree the loss budget before the incident

basics

~20 s

It is a data-loss decision, not a coding one. Stall only if the upstream can absorb waiting without failing itself; shed only for data whose owner has agreed it may be lost, per class, with a counter and an alert. Decide it before the incident and make it configurable.

solid answer

~50 s

Start from who owns the data and what the source does when you stop reading. If the upstream is a durable log or a producer that retries, stalling is nearly free — the events wait somewhere that can hold them, and you have exported your saturation honestly. If the source is a live stream with no replay, or a caller whose own deadlines will fire, stalling converts your slowness into their outage and you should shed instead. That means shedding is only defensible for data whose owner has agreed to lose it, so I split traffic by class: billing-grade events stall or spill to durable storage, telemetry gets sampled away first. Whatever the choice, the shed path must be counted, exported and alerted on, and the thresholds must live in configuration so the policy can change during an incident without a redeploy. And I write it down: the data's owner can overrule me, and that conversation belongs before the flash sale, not during it.

go deeper

for a junior

Understand that a full intake queue forces a choice between making the producer wait and throwing work away, and that the code cannot decide which is correct on its own.

for a middle

Be able to argue both sides concretely: what stalling does to the upstream, what shedding costs downstream, and why the shed path needs a counter so the loss is visible.

for a senior

Show that you would classify traffic and shed the cheapest class first, define what deserves a page, and validate the policy with a load test that actually pushes past capacity.

for a principal

Own the loss budget as a commitment negotiated with the data's owner, keep the levers in configuration so the policy is changeable during an incident, and be explicit that durable spill is the priced alternative when neither loss nor delay is acceptable.

## Why this is not a coding decision Once the intake is a bounded channel, the code has exactly one open question: what happens when the send would block. Blocking means the producer waits and the stall travels upstream. Giving up means an event ceases to exist. Both are correct implementations; which is right depends on facts the code cannot see — whether the data is replayable, who is harmed by loss, and who is harmed by delay. That is why this is an ownership question. The service owner proposes; the owner of the data can overrule. ## The first question: what does the source do when you stop reading? Stalling is only *free* if the pressure lands somewhere that can hold it. - **A durable upstream** — a log or queue with retention, a producer that retries with its own buffer — absorbs the stall. Consumption falls behind, retention covers the gap, and the visible symptom is lag rather than loss. Stall. - **A live, non-replayable stream** — metrics, traces, click events, anything with no acknowledgement — discards what you do not take, whether or not you admit it. Here "stalling" is really "dropping, invisibly, somewhere else". Shed explicitly instead, so at least the loss is counted. - **A synchronous caller with its own deadline** — an HTTP or RPC client that will time out in two seconds — turns your stall into their failure. Worse, if that caller holds a connection or a pool slot while waiting, your slowness spreads into their capacity and you have manufactured a cascade. Shed fast, and answer with a clear rejection so they can retry or degrade. ## The second question: what is the data worth? A single global policy is almost always wrong because a service rarely carries one kind of event. Classify: - **Must not be lost** — financial events, audit records, anything that has to reconcile. These stall, and if stalling is unacceptable too, they need a durable landing zone: spill to disk or push back to a persistent queue. That costs money and complexity, and it is the honest answer when neither loss nor delay is acceptable. - **Lossy by nature** — telemetry, sampled traces, debug events. These are shed first, and preferably sampled rather than cliff-dropped, so the signal degrades in quality instead of vanishing. - **Everything else** — decide deliberately and write it down, because the default in an incident will be whatever the code happens to do. A good design sheds in that order automatically, so the first thing lost under pressure is the cheapest thing. ## The third question: what do you page on? Depth pinned at capacity is a saturation signal, not necessarily a problem — a burst absorbed and drained is the system working. What deserves a page is sustained saturation, queueing delay past the freshness the consumers of this data were promised, or a shed rate above the agreed budget. Committing to that number is the real deliverable: "we will lose at most X% of class-B events over an hour, and you will be told when we do" is a statement a data owner can accept or reject. "The channel has 1024 slots" is not. ## What makes the decision reversible Everything about this changes under load, so the levers must be operable under load: - Timeouts, queue capacity, sampling rates and the per-class policy in configuration, not constants. - The shed path instrumented from the first line of code, with a counter per class, so the cost of the policy is measured rather than assumed. - A load test that actually drives the service past capacity, because the policy is untested until something has been dropped in anger. - A written note of who agreed to what. The uncomfortable version of this incident is the one where the on-call engineer discovers, at 3am, that they are the person choosing which customer's events to discard. ## The trap to avoid The most common failure is a third option that looks like neither: enlarging the buffer. It feels like a compromise — no stalling, no dropping — but it only converts the decision into queueing delay and memory, and it postpones the moment of choice to a worse time, with a bigger backlog of stale events that may no longer be worth processing when they finally come out. Buffer sizing absorbs bursts; it does not resolve a sustained rate gap. If the arrival rate exceeds the drain rate for long enough, something waits or something is lost, and the only question a leader controls is whether that was chosen deliberately in advance or discovered during an outage.

  • Your upstream is a synchronous caller with a two-second deadline. Does that change the answer?
    Yes — it removes stalling as an option beyond a fraction of that budget. Waiting past their deadline gives them a timeout instead of an answer while your goroutine still holds the work, and if they hold a connection or pool slot meanwhile, your saturation consumes their capacity. Reject quickly with a distinguishable error so they can retry or degrade on their own terms.
  • Under sustained overload, why is enlarging the buffer not a third option?
    It converts the choice into queueing delay and memory rather than resolving it. The rate gap still integrates, you reach the same wall later with a larger backlog, and the events that finally emerge may be too stale to be useful. Buffers absorb bursts; only stalling or shedding closes a sustained gap.
  • How do you make this policy operable at 3am by someone who did not design it?
    Put the capacity, timeouts, sampling rates and per-class policy in configuration so they can be changed without a redeploy. Export a shed counter per class and a queueing-delay metric so the cost of each setting is visible. And write down which classes may be dropped and who agreed, so the on-call engineer is executing a decision rather than making one.
  • What if neither loss nor delay is acceptable for a class of events?
    Then the events need somewhere durable to wait — spilling to local disk or pushing back to a persistent queue — and that is a real cost in complexity, storage and operations. Say so plainly and price it, rather than pretending a larger in-memory buffer provides the same guarantee. Unbounded memory is not durability.

A restaurant kitchen behind on tickets can stop taking orders at the door or bin the least important ones. Either is workable; what is not workable is stacking tickets on the counter until the kitchen collapses — and neither choice belongs to the person holding the tickets at midnight.

saying these in an interview costs you the question

  • Picks one global policy for every kind of event
  • Treats a bigger buffer as a way to avoid the decision
  • Sheds silently with no counter, alert or owner sign-off
  • Stalls a caller that has its own short deadline
  • Leaves the choice to whoever is on call during the incident
  • Confuses an in-memory buffer with durable storage