skip to content

In a Pipes and Filters pipeline, what's the difference between a push-based and a pull-based data flow model, and what trade-off does each make?

level: middleimportance: must knowfreq 55%

answer

  1. push = producer-driven, consumer reacts
  2. pull = consumer-driven, calls next()
  3. push risks unbounded buffering
  4. pull stalls on a blocked consumer
  5. Reactive Streams request(n) = pull+push hybrid

basics

~20 s

Push means the upstream filter sends data downstream as soon as it's ready, like shoving items down a conveyor belt. Pull means the downstream filter asks for the next item when it's ready to handle it, like taking items off a shelf one at a time.

solid answer

~50 s

In push-based dataflow, each filter actively writes its output to the pipe as soon as it has produced it, and the downstream filter reacts whenever data arrives — control flow is driven by the producer. In pull-based dataflow, each filter requests the next unit of data from its upstream pipe only when it's ready to process it (think an iterator's next() call) — control flow is driven by the consumer. Push suits scenarios with unpredictable or continuous arrival (sensor streams, UI events) since producers don't have to wait for a consumer poll, but it needs some buffering/backpressure mechanism or a fast producer can overwhelm a slow consumer. Pull naturally paces itself to the slowest consumer, so it needs no separate backpressure mechanism, but it doesn't fit sources that generate data on their own schedule, and it makes fan-out to multiple independent consumers awkward since each would need to pull independently.

go deeper

for a junior

Should be able to describe push as 'sender decides when' and pull as 'receiver decides when' with a simple example of each.

for a middle

Should explain the buffering/backlog risk under push and the stalling risk under pull, and identify which fits a given scenario.

for a senior

Should discuss hybrid demand-signaling approaches (e.g. request(n)-style protocols) and design a bounded-buffer bridge between mismatched push/pull stages.

for a principal

Should reason about fan-out and multi-consumer scenarios, and make an informed style choice (push, pull, or hybrid) as part of an overall pipeline architecture given source characteristics and SLAs.

## Two ways control can flow Push and pull are the two fundamental ways control can flow through a pipe connecting two filters, and the choice between them shapes almost everything about how a pipeline handles speed mismatches, memory, and failure — it's one of the first concrete design decisions you make once you move past the abstract idea of 'filters connected by pipes.' ## Who is in control - **Push.** In a push model, the producer filter is in control: as soon as it has a unit of output ready, it writes it to the pipe and moves on, without waiting for the consumer to ask for it. A clean push example is an event-driven filter that registers a callback and gets invoked by whatever upstream mechanism produces records, with no ability to say 'don't call me yet.' - **Pull.** In a pull model, the consumer filter is in control: it calls something like `next()` on its upstream source and blocks (or gets nothing back) until the upstream filter has a unit ready, only advancing when it is itself ready for more work. An iterator, a generator, and a database cursor are the classic pull abstraction — the consumer literally drives the pace of the entire chain by how often it asks for more. | | Push | Pull | |---|---|---| | **In control** | the producer filter | the consumer filter | | **Concrete form** | an event-driven filter that registers a callback | an iterator, a generator, a database cursor | | **Natural fit** | sensor readings, incoming network packets, UI click events | a batch job reading rows from a database cursor | | **Failure mode** | unbounded queue growth, OOM kills, silent data loss | a slow or blocked consumer stalls the entire chain upstream of it | ## Which one fits a source The reason this distinction exists is that real filters don't all run at the same speed, and someone has to decide what happens when they don't. - **Push** is a natural fit when the producer's timing is externally determined and can't be deferred — sensor readings, incoming network packets, UI click events — the data simply arrives when it arrives, and the pipeline's job is to keep up. - **Pull** is a natural fit when you want the consumer to set the pace and you don't want unconsumed data piling up anywhere — a batch job reading rows from a database cursor typically pulls, because there's no inherent urgency to have the next row ready before anyone asks for it. ## The trade-off The trade-off is essentially about who bears the cost of a speed mismatch and how visible that cost is. With pure push, if the consumer is slower than the producer, output has nowhere to go except an ever-growing buffer, which either grows unbounded (risking out-of-memory) or is capped, in which case you need an explicit backpressure signal — without that, you silently drop data or crash. Pull avoids this by construction: the consumer only ever asks for the next item when ready, so there's no unconsumed backlog to manage, but the cost shows up elsewhere — a producer that generates data on its own schedule can't simply be paused between pull requests, so pure pull is a poor fit for genuinely push-driven sources, and you often need an adapter (a bounded buffer) to bridge the two, which reintroduces the same backlog question in a smaller, more controlled form. Pull also complicates fan-out: if two independent downstream filters each want to pull from the same upstream filter at their own pace, you either need the upstream to support multiple independent cursors or you accept that one slow puller stalls the shared source for everyone. ## Failure modes in production In production, each model has its own failure mode. - **The push failure mode** shows up as unbounded queue growth leading to OOM kills, or as silent data loss when a bounded queue drops the oldest/newest item under pressure — this is one of the most common causes of a streaming pipeline falling over under load spikes it handled fine at steady state. - **The pull failure mode** shows up differently: a slow or blocked consumer stalls the entire chain upstream of it, which can starve unrelated consumers sharing the same source, or cause request timeouts in systems where the 'pull' is implemented as a synchronous call with a deadline. ## Where the two meet A concrete real-world instance of a hybrid approach is the Reactive Streams specification (implemented by libraries like Project Reactor, RxJava, and Akka Streams), which explicitly models a pull-with-push-notification protocol: the consumer signals how many items it's willing to receive via a `request(n)` call, and the producer pushes up to that many items and then waits — a demand-signaling design created specifically because pure push causes unbounded buffering and pure pull can't handle sources that genuinely produce data on their own schedule, like network sockets.

  • If a pipeline stage genuinely can't be paused (e.g. a live audio capture device), how do you combine that with a pull-based downstream consumer that processes more slowly?
    You insert a bounded buffer (a ring buffer or queue) between the push-driven source and the pull-driven consumer, so the source keeps pushing into the buffer at its own pace while the consumer pulls from the buffer at its own pace. You still need a policy for what happens when the buffer fills — block the producer, drop the oldest samples, or drop the newest — and that policy choice is itself a real design decision with audible/visible consequences for an audio pipeline specifically.
  • Why doesn't pure pull dataflow need an explicit backpressure mechanism the way push does?
    Because control flow already originates from the consumer — it only asks for more data when it's ready for more, so by construction there's never a batch of unconsumed data building up anywhere in the chain. Backpressure is really a mechanism to retrofit that same 'don't send me more than I can handle' property onto a push model, so pull gets it for free at the cost of not fitting sources that can't be told to wait.

Push is a chef plating dishes and sending them out the moment they're ready, piling up on the pass if the servers can't keep up; pull is a buffet where diners walk up and take food only when they're ready to eat, so nothing piles up but the kitchen still has to keep the buffet stocked.

saying these in an interview costs you the question

  • Says push and pull are interchangeable with no trade-off
  • Doesn't mention what happens to unconsumed data under push when producer outruns consumer
  • Claims pull has no downsides
  • Can't name a concrete pull abstraction (iterator/generator/cursor) or push abstraction (callback/event)
  • Thinks backpressure is only relevant to push models without explaining why

context