skip to content

Backpressure

What happens when a producer outruns its consumer: demand signalling that makes flow control explicit, and the buffer, drop, latest and fail strategies. Interviewers use it to test real understanding.

on this pageshow

questions

18

In a log-shipping pipeline whose hold between reader and slow archival writer has no capacity limit, what fails and when?

level: middleimportance: must knowfreq 66%

answer

  1. a mismatch, not a spike
  2. the queue lives in memory
  3. arrival rate minus drain rate
  4. deficit times elapsed time
  5. a limit turns growth into a decision

basics

~20 s

An unbounded hold converts a sustained rate mismatch into memory exhaustion. Every line the writer cannot take is retained, so the pipeline behaves normally for as long as the spare memory lasts, then the process dies all at once.

solid answer

~50 s

A hold with no capacity limit does not fix a slow consumer; it stores the difference between the two rates. If the reader is faster than the writer only during a burst, the hold drains afterwards and nothing is wrong. If the reader is faster on average, there is no equilibrium: retained items grow at `arrival rate - drain rate`, and the pipeline keeps running until the process cannot allocate. That is why the outage arrives hours after the mismatch started and looks sudden — the deficit is small, the spare memory is large, and no stage reports anything, because retaining items is exactly what the hold was told to do. Bounding the capacity does not make the writer faster; it converts a silent memory failure into a decision that fires early, at the boundary, where it can be seen.

code

pseudocode · 7 lines
pseudocode
pipeline = readLogLines()                       // about 12000 lines per second
    .holdInMemory(capacity = UNLIMITED)         // the queue's real limit is the heap
    .writeEach(line -> archive.append(line))    // drains about 11500 lines per second

// retained(t) = (12000 - 11500) * t  =  500 * t lines
// at about 400 bytes per retained line: 200 KB per second
// with 3 GB spare: 15000 seconds, roughly 4.2 hours to allocation failure

go deeper

for a junior

Remember that unconsumed values have to live somewhere, and that somewhere is process memory. A stream with a fast source and a slow destination is holding the difference right now.

for a middle

Be able to state the invariant and do the arithmetic: retained items grow at arrival rate minus drain rate, so a tiny percentage shortfall becomes gigabytes over hours. Separate the bounded burst from the sustained mismatch.

for a senior

Show that you would have caught it: depth exported as a metric, an alert on depth that never returns to baseline, and a capacity limit declared everywhere a stage can retain so the failure is early and local rather than an exhausted process.

for a principal

The judgment is where the failure should land. Unbounded retention moves a throughput problem into the memory subsystem, where it takes the whole process with it; bounding it makes the pipeline fail earlier and more often, and someone has to be ready to act on that new signal.

## What an unbounded hold actually promises Between a fast producer and a slow consumer, a stream pipeline has to put the unconsumed values somewhere. That somewhere is a **hold**: an in-memory queue owned by one stage, keeping items the next stage has not asked for yet. When the hold is declared with no capacity limit, the promise it makes is "I will never refuse an item". It is easy to read that as "the pipeline can now cope with a slow writer". It is not what it says. It says the queue's real limit is whatever memory the process has, and that limit is discovered by hitting it. The important consequence: **a hold does not change any rate**. The reader still reads at its rate, the archival writer still drains at its rate. The hold only decides where the difference between them accumulates. ## A burst and a mismatch look identical for the first minute The two cases behave the same at the start and end completely differently. | | Bounded burst | Sustained mismatch | |---|---|---| | Cause | Input spikes above drain rate for a while | Input average exceeds drain average | | Depth over time | Rises, then returns to near zero | Rises, never returns to baseline | | Peak retained | Bounded by burst size | Bounded only by memory | | Correct design | Capacity >= expected peak | No capacity fixes it; the consumer or the input must change | | Failure mode | None, if the peak fits | Allocation failure, hours later | This is why an unbounded hold survives staging and dies in production. A test rig replays a finite file: the input ends, the hold drains, the run is green. Production never ends. ## Doing the arithmetic Suppose the reader tails application logs at **12,000 lines per second** and the archival writer sustains **11,500 lines per second**. The deficit is 500 lines per second — a shortfall of about 4%, far too small to notice on a dashboard of throughput. - Each line, with its parsed fields and object overhead, costs roughly **400 bytes** retained. - Retained bytes grow at `500 x 400 = 200 KB per second`, which is about **720 MB per hour**. - With **3 GB** of spare memory, the process runs for about `3 GB / 200 KB per second = 15,000 seconds`, roughly **4.2 hours**, and then fails. Run it again after a restart and it fails at the same slope after the same 4 hours, because nothing about the rates changed. n here is the retained item count, and it is linear in elapsed time, not in traffic bursts. ## Why nothing warns you Every stage is behaving to contract: - The reader is emitting values it was asked to emit. - The hold is retaining values, which is its entire job. - The writer is completing every write it starts, just slowly. - No error signal travels the stream, because no failure has happened yet. The only observable that moves is the hold's depth, and the depth of an unbounded hold is usually not exported anywhere. The first alert is the allocation failure, at which point the pipeline also loses every line it was holding. ## What a capacity limit changes A bounded hold does not add throughput. It changes three things: 1. **The failure becomes a decision.** Someone chooses, at design time, what happens when the hold is full. That choice belongs to the pipeline's owner rather than to the allocator. 2. **The failure arrives early.** The limit is reached minutes into a mismatch instead of hours, while the backlog is still small and the cause is still visible. 3. **The blast radius shrinks.** A full hold affects one pipeline stage; an exhausted process takes down everything else sharing it, including the part of the service that could have reported the problem. Bounding is therefore the default position: declare a capacity everywhere a stage can retain, and treat "unlimited" as a claim that needs evidence, not as a safe starting point. ## What an interviewer is listening for The weak answer is "add a buffer so the fast producer does not overwhelm the slow consumer", stated as if buffering were a solution. The strong answer names the invariant — a buffer stores a rate difference and cannot change a rate — then separates the bounded burst (a legitimate use) from the sustained mismatch (where only a slower producer, a faster consumer, or a deliberate loss policy helps), and finishes with the operational point: an unbounded hold hides a throughput problem until it is a memory outage, and memory is the worst place to discover it.

  • The reader outruns the writer only during a nightly burst. Is an unbounded hold safe there?
    Safer, but still a bet. A bounded burst drains afterwards, so the risk is only the peak depth times the item size. The bet is that the burst size is known and stays known. Declaring a capacity at that peak costs nothing and turns a bad night into an explicit event instead of an allocation failure.
  • Why does no error surface while the hold is growing?
    Because nothing has gone wrong yet by any stage's contract. Retention is the behaviour the hold was configured for, the writer completes every write it starts, and no failure signal is generated. The only moving observable is depth, which an unbounded hold usually does not export. The first report comes from the allocator.
  • Does moving the hold to disk instead of memory solve it?
    It buys time proportional to the extra capacity and nothing else. A sustained deficit fills any finite store, and the added depth is also added waiting time for every item. It converts a fast failure into a slow one plus a freshness problem, so it helps only when the mismatch is genuinely temporary.

A tap running slightly faster than the drain does not make the basin overflow later or maybe — the overflow is certain from the moment the rates differ. The basin's size only decides what time it happens.

saying these in an interview costs you the question

  • Buffering fixes a slow consumer rather than storing the difference
  • A large enough hold removes the need for any capacity limit
  • Steady memory growth in a pipeline must mean an object leak
  • The problem would show up immediately, so a short test would catch it
  • Depth only matters because of memory, so an idle machine is safe
open as a page

In a demand-driven stream, a consumer requests 10 values, then 5 more before any arrive — what may the producer send?

level: middleimportance: must knowfreq 68%

basics

~20 s

Up to 15 values, and not one more. Requests accumulate additively into a single outstanding count, each delivered value spends one unit, and the producer is barred from sending a sixteenth value until the consumer asks again.

open as a page

A vehicle-position feed and a settlement-instruction feed both outrun their consumers - why can one discard values and the other not?

level: middleimportance: must knowfreq 58%

basics

~20 s

Classify the item first. A position is a snapshot that the next one supersedes, so discarding intermediates costs the consumer nothing. A settlement instruction has an effect of its own that no later item re-derives, so discarding it is a lost transfer.

open as a page

A bounded buffer in a stream pipeline fills up - what do drop-newest, keep-latest and fail-fast each sacrifice?

level: middleimportance: must knowfreq 66%

basics

~20 s

Bounded buffering spends memory and adds latency and only postpones the decision; dropping the newest arrival sacrifices freshness; keeping only the latest sacrifices every superseded value; failing fast sacrifices availability but is the only one that reports the loss.

open as a page

A fixed-interval clock source emits a tick whether or not the consumer asked for one — where must flow control live instead?

level: middleimportance: must knowfreq 60%

basics

~20 s

Flow control moves into the adapter wrapping the source. A clock cannot wait for demand, so the adapter decides each tick's fate: discard it, overwrite a stored latest value, or hold it in a bounded buffer.

open as a page

Your log-shipping service's memory climbs steadily for hours and then dies; how do you show it is a rate mismatch, not a leak?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Compare rates and run a drain test. A rate mismatch retains items that belong in a queue and recovers when the input drops below the drain rate; a leak retains objects nothing needs and recovers only on restart.

open as a page

Turnstile scans arrive faster than the counting consumer can take them — how do you choose the adapter's boundary policy?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Let the consumer's need decide. If every event must be counted, hold a bounded buffer and fail when it fills, or aggregate in the adapter; if only the newest value matters, overwrite; if a steady cadence is enough, release the latest value on a clock.

open as a page

A deep hold keeps a log-shipping pipeline alive, but archived lines now arrive nearly an hour late — why?

level: middleimportance: should knowfreq 52%

basics

~20 s

Queue depth is waiting time. By Little's Law a steady depth divided by the drain rate is how long every arriving item waits, so a hold deep enough to survive a mismatch also delays each line by exactly that much, in order.

open as a page

Which stages in a log-shipping pipeline retain items internally even though nobody declared a buffer?

level: middleimportance: should knowfreq 44%

basics

~20 s

Any stage that cannot emit until it has seen more than one item retains them. Time-based grouping holds a whole window, concurrent flattening holds every live inner source and its results, and handing values to another worker needs a queue.

open as a page

A stage in an export pipeline fetches ledger rows a page at a time but emits them one by one — how should it convert downstream demand into upstream requests?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Not one for one. The stage keeps two counters — what it owes downstream and what it has asked upstream — requests a whole page ahead, and replenishes upstream only when the page has drained to a low-water mark, never asking for more than it has room to hold.

open as a page

A broadcast hub pushes each room message to thousands of WebSocket subscribers, and one slow subscriber delays everyone - where should the buffer sit, and why?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Per subscriber, not shared: encode each message once, put a reference on each subscriber's own bounded queue, and drain each queue with its own writer. Overflow then drops, keeps only the latest, or disconnects only the laggard.

open as a page

A stream quietly discards values whenever load spikes and every dashboard still looks healthy - how do you make that loss visible?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Count every discard at the point it happens, publish it as a rate beside throughput, and carry a gap marker with the next surviving value so consumers can see what they missed. Escalate sustained loss into a failure rather than leaving it silent.

open as a page

An adapter over a push feed accepts demand but emits every value regardless — what does that break?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A dishonest adapter breaks the guarantee the rest of the pipeline is sized against: downstream stages assume no more than outstanding demand arrives. The excess then lands wherever the pipeline is slowest, so the symptom appears far from its cause.

open as a page

Two teams' stream pipelines have now died from silent buffering; what would you require of every pipeline before it ships?

level: principalimportance: should knowfreq 38%

basics

~20 s

Require four things: no unbounded retention anywhere, a capacity derived from the freshness budget and checked against memory, depth and delivery age exported per retaining stage, and a soak at a deliberate deficit proving memory stays flat.

open as a page

Where would you draw the line for permitting unbounded demand in your organisation's stream pipelines?

level: principalimportance: should knowfreq 36%

basics

~20 s

Default to refusing it, and permit it only where something other than demand already bounds the flow: a source finite by construction, or a consumer that does its work inline so the serialised handover paces the producer. Require the bound to be named, not assumed.

open as a page

Across many streaming services each team picks its own overflow policy by habit - what would you standardise, and what stays a team decision?

level: principalimportance: should knowfreq 34%

basics

~20 s

Standardise the decision procedure, not the policy: every stream that may discard must classify what one item means, record the policy it chose and why, and publish a count of what it discards. The policy itself and the buffer size stay with the team.

open as a page

Several teams each wrap a push feed that cannot be slowed — what standard do you set for their adapters?

level: principalimportance: should knowfreq 34%

basics

~20 s

Standardise the invariants, not the policy: no adapter exceeds outstanding demand, every loss policy is declared rather than emergent, discards are counted and exported, and no hold is unbounded. Which policy each feed uses stays with the team that owns its consumer.

open as a page

A consumer's outstanding demand for a ledger stream is 4 — may the producer deliver those four values concurrently?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

No. Demand is permission to send a number of values, never permission to send them at once. Delivery to one subscriber is serialised: each value is handed over only after the previous handover has returned.

open as a page