skip to content

questions

4

In a reactive-streams style pipeline the producer pushes items to the consumer, yet the consumer sends demand signals upstream — messages that say request N more items. If delivery is push-based, why is that demand channel needed at all?

level: middleimportance: must knowfreq 45%

answer

  1. Push delivery, pull volume
  2. Never emit beyond outstanding demand
  3. Unbounded push = unbounded queue
  4. TCP receive window at app level
  5. Batch requests; unbounded request opts out

basics

~20 s

Pure push has no rate control: a fast producer and slow consumer means unbounded queue growth and eventual memory exhaustion. Demand signalling lets the consumer declare capacity, and the producer may never emit more than the outstanding demand, so items in flight stay bounded without blocking any thread.

solid answer

~60 s

Push delivery is efficient — no thread parks waiting, and items arrive as soon as they exist — but it puts the producer in charge of the rate. If the producer is faster than the consumer, the excess must go somewhere: an unbounded queue (memory exhaustion), a blocked thread (the cost push was meant to avoid), or the floor (data loss). Demand signalling splits the two decisions. **Delivery stays push**: the producer sends when it has something. **Volume becomes pull**: the consumer sends demand upstream, and the protocol forbids emitting more items than the outstanding demand. In-flight items are therefore bounded by what the consumer asked for, and nothing blocks — the producer simply stops emitting until more demand arrives. The effect is a pipeline that self-limits at the rate of its slowest stage, with bounded memory. It is the same idea as a receive window in TCP: the receiver advertises how much it can take, and the sender must respect it. Requesting in batches keeps the signalling overhead small. What the producer should do when it cannot slow down is a separate design question.

code

text · 8 lines
text
consumer -> producer : request(4)
producer -> consumer : item, item, item, item
producer            : (demand now 0 -> stops emitting, no thread parked)
consumer            : processes two, then
consumer -> producer : request(2)
producer -> consumer : item, item
...
producer -> consumer : complete    # terminal, nothing after it

go deeper

for a junior

Say that a fast producer plus a slow consumer means an ever-growing queue, and demand signalling stops the producer from sending more than the consumer asked for.

for a middle

Separate the two axes — push delivery, pull volume — and explain that in-flight items are bounded by construction with no thread parked.

for a senior

Bring in the TCP receive-window parallel, batching of demand, the unbounded-request escape hatch, and the fact that the slowest stage ends up throttling the source.

for a principal

Frame it as graceful degradation: the system slows instead of exhausting memory, and the residual risk lives at sources whose rate you do not control.

## Two independent axes: who delivers, who sets the rate It helps to see that a data transfer between two asynchronous parties has two separable decisions. - **Who initiates delivery?** In pull, the consumer asks and waits. In push, the producer sends when the item exists. - **Who controls the volume?** Whoever decides how many items may be outstanding. Classic pull couples both to the consumer, which gives free flow control but usually parks a thread per stream. Classic push couples both to the producer, which is thread-efficient and low-latency but gives the consumer no way to say *slow down*. Reactive streams split the axes: **push delivery, pull volume**. The consumer sends a demand signal — conceptually *request(n)* — and the producer is contractually forbidden from emitting more than the sum of outstanding demand. When demand hits zero, emission stops; when the consumer finishes processing and asks for more, emission resumes. ## What goes wrong without demand With unrestrained push and a producer faster than the consumer, the surplus has three possible destinations, none of them free: 1. **A queue.** If it is unbounded, memory grows without limit until the process dies — usually after a long period of rising latency and garbage-collection pressure that makes the failure look like something else. This is the classic incident: the system runs fine for hours and then falls over under a traffic spike, with no single slow component to blame. 2. **A blocked thread.** Making the producer block when the buffer is full does bound memory, but reintroduces exactly the thread cost that motivated the async model, and can deadlock if the producer's thread is also the one that drains the queue. 3. **The floor.** Dropping is legitimate for some data, catastrophic for the rest, and either way it is a policy decision that should be explicit rather than emergent. Demand signalling removes the surplus at the source: an item that would exceed demand is simply not emitted. ## Why the flow-control analogy is exact This is the receive window of TCP moved into application-level pipelines. The receiver advertises capacity; the sender may have at most that much unacknowledged data in flight. Nobody blocks, nothing is dropped, and the transfer runs at the rate of the slower party. Reactive streams generalize it: the *slowest stage anywhere in the pipeline* eventually throttles the source, because each stage only requests upstream what it can pass downstream. ## The protocol rules that make it work The demand contract is only half of the model. The other half is a small set of signalling rules that let independent operators, written by different people, be composed safely: - **A bounded item signal, plus terminal signals.** A stream emits zero or more items, then at most one terminal signal: completion or failure. After a terminal signal, nothing more may be emitted. - **Serial delivery.** Signals to a given consumer are delivered one at a time, never concurrently, so operator implementations do not need internal locking for the delivery path. This is why an async pipeline can be assembled from stages that are individually single-threaded in their logic while the pipeline as a whole is concurrent. - **Demand is cumulative and non-negative.** Requests add up; the producer tracks the outstanding total. - **Cancellation.** The consumer can withdraw interest at any time, which stops the source and releases resources. ## Practical consequences to mention - **Batch your demand.** Requesting one item at a time makes a signal round trip per item and can dominate cost. Real implementations request in batches and replenish when a fraction is consumed, so the pipe stays full while staying bounded. - **Requesting unbounded opts out.** A consumer may request an effectively infinite amount, which switches the stage back to unrestrained push. That is legitimate for genuinely fast consumers and a common accidental cause of memory growth. - **Demand only helps if the source can slow down.** A source that is itself pull-capable — a database cursor, a file, a paginated API — can honour demand exactly. A source with an external clock — sensor readings, user input, market ticks — cannot be told to wait, so the pipeline must adopt a policy for the excess. Choosing among those policies is a deep-dive of its own and sits outside this model-level question. - **The prize is behaviour under overload.** The system slows down rather than falling over, which is the property you actually want during an incident. ## In an interview Say: push delivery is thread-efficient but hands rate control to the producer; demand signalling returns rate control to the consumer without parking anyone, so items in flight are bounded by construction; it is TCP's receive window at application level; batch the requests; and it only works when the source can be slowed.

  • Can two item signals be delivered to the same consumer concurrently?
    No. The protocol requires signals to a given consumer to be serial — one at a time, never overlapping — and at most one terminal signal, after which nothing more may be sent. That rule is what lets operator authors write their delivery logic without internal locking, and it is what makes independently written stages composable across thread boundaries.
  • What breaks when a consumer requests an unbounded amount of demand?
    The stage reverts to unrestrained push: the producer may emit as fast as it can, and any speed mismatch accumulates in a queue somewhere. That is fine when the consumer is genuinely faster than the source, and it is a frequent accidental cause of memory growth otherwise, because the pipeline looks reactive but has silently opted out of flow control.

A waiter who brings dishes the moment they are ready will bury a slow eater. Demand signalling is the diner saying bring two more when I am ready: the kitchen still delivers as soon as it can, but never more than two plates sit on the table.

saying these in an interview costs you the question

  • Calling demand signalling a form of blocking, when its purpose is bounded flow without parking a thread
  • Believing an unbounded in-memory queue is a solution to a rate mismatch rather than a delayed failure
  • Assuming demand can throttle any source, including ones driven by an external clock that cannot be slowed
  • Requesting one item at a time and treating the resulting signal overhead as inherent to the model
  • Thinking signals may be delivered concurrently to a consumer, so every operator needs its own locking

context

open as a page

You already have futures and promises for asynchronous results. When is a future the wrong abstraction and an asynchronous stream the right one?

level: juniorimportance: should knowfreq 38%

basics

~20 s

A future carries exactly one outcome, once. A stream carries zero to many items over time plus a terminal completion or failure. Use a stream when results arrive incrementally, when there may be many or infinitely many, or when the producer's rate could outpace the consumer.

open as a page

Compare three ways for a consumer to receive a sequence of items from a producer: a blocking pull loop where the consumer asks for the next item, unrestrained push where the producer invokes a consumer callback, and demand-driven push where the consumer sends batched requests upstream. What does each cost, and when does the third win?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Blocking pull gives free rate control but parks a thread per stream. Unrestrained push parks nothing but hands rate control to the producer, so surplus items pile up. Demand-driven push combines them: batched permits bound memory, push delivery keeps threads free. It wins with many concurrent, remote, rate-mismatched sources.

open as a page

A team is deciding whether to build a new service's asynchronous layer around demand-driven reactive streams, around suspending sequential code (coroutine or async-await style), or around plain futures dispatched on an event loop. How would you reason about that choice?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Choose by problem shape. One result per request means futures or suspending code. Many items over time with a real rate mismatch and a bounded-memory requirement justifies streams. Weigh debuggability, the viral spread of the model through signatures, and whether the whole path is genuinely non-blocking.

open as a page