skip to content

Some messaging systems have the broker push messages to consumers, while others have consumers pull (poll) messages from the broker. How do these two models differ in how they handle a consumer that's overwhelmed, and why do most high-throughput systems favor pull?

level: seniorimportance: should knowfreq 55%

answer

  1. push: broker decides pace; pull: consumer decides pace
  2. pull back-pressure is implicit/free
  3. push needs bolted-on flow control (e.g., prefetch/QoS credit)
  4. pull trade-off: polling latency and overhead
  5. long-polling as a hybrid

basics

~20 s

Push means the broker decides when to send you the next message, which risks flooding a slow consumer. Pull means the consumer asks for more only when it's ready, so it naturally never asks for more than it can handle.

solid answer

~60 s

In a push model, the broker actively sends messages to the consumer as they arrive or become available, so the broker controls the rate — a consumer that can't keep up either gets overwhelmed (unbounded in-flight messages, memory pressure, dropped connections) or the protocol needs an explicit flow-control extension (like TCP-style windowing or an explicit 'pause/resume' signal) bolted on to prevent that. In a pull model, the consumer initiates each fetch, requesting the next batch only when it's ready for more, so back-pressure is implicit and free: a slow consumer simply pulls less often, and the broker never sends unrequested work. Most high-throughput log-style brokers (e.g., Kafka) are pull-based specifically because it lets each consumer control its own pace without any special-cased flow-control protocol, and because polling in batches amortizes network round-trips efficiently. The trade-off is that pure pull adds latency (a consumer only learns about a new message on its next poll, not the instant it's produced) and involves polling overhead when idle; some systems use long-polling or hybrid push-pull to get low latency for the push side while keeping pull-model back-pressure.

go deeper

for a junior

Should grasp the basic direction of control: push means the broker sends, pull means the consumer asks.

for a middle

Should explain that pull naturally limits how much a consumer receives, while push needs some extra mechanism to avoid overwhelming a slow consumer.

for a senior

Should discuss concrete mechanisms — prefetch/QoS counts, long-polling — and the latency-vs-back-pressure trade-off between the two models.

for a principal

Should reason about why a given system (e.g., Kafka) chose pull for its scalability profile, and evaluate when push's lower latency is worth its added flow-control complexity for a specific workload.

## Who initiates each message The mechanism difference is about who initiates transfer of each message and therefore who controls its pace. | Model | How the transfer works | |---|---| | **Push** | In a push model, the broker maintains an open connection or subscription to each consumer and actively sends messages as soon as they're available (or as soon as the broker decides to send them), without the consumer having to ask. Classic examples include traditional AMQP-style brokers (like RabbitMQ) in their default consume mode, and webhook-style delivery where the sender initiates an HTTP POST to the receiver. | | **Pull** | In a pull model, the consumer periodically calls a fetch/poll operation, explicitly requesting the next chunk of messages, and the broker only responds to that specific request — it never sends anything the consumer didn't ask for. Kafka is the canonical pull-based example: consumers repeatedly call a `poll()`-style method, specifying how many messages or how much data they're willing to receive, and the broker returns up to that amount from the offset the consumer is at. | ## Where rate-control lives The reason this distinction exists, and matters enormously, is that it determines where rate-control naturally lives. - **In push**, the sender (broker) is the one deciding the rate, but the sender generally doesn't know the receiver's real-time capacity — it knows the message is ready to go, not whether the consumer's CPU, memory, or downstream dependency can absorb it right now. That mismatch is exactly the back-pressure problem from before, but now happening at the transport/protocol layer instead of just the application/business layer: if push keeps firing messages at a consumer that's falling behind, messages queue up somewhere (in-memory buffers on the consumer side, unacknowledged in-flight counters on the broker side, TCP socket buffers) with the same risks — memory exhaustion, dropped connections, or unbounded unacknowledged message counts. - **In pull**, the consumer is structurally the one deciding the rate: it only issues the next poll when it's ready, so a consumer that's overwhelmed simply doesn't call poll again until it's caught up, and the broker is never in a position to force-feed it more than requested. This is why pull-based back-pressure is often described as "free" or "implicit" — it falls directly out of the request/response shape of the protocol rather than needing a bolted-on flow-control extension. ## The trade-off runs both ways The trade-off is not one-directional, though. - Push systems that want to avoid overwhelming consumers typically add explicit **flow-control mechanisms** to compensate — AMQP's QoS/prefetch-count setting caps how many unacknowledged messages the broker will have in flight to a given consumer at once, effectively converting push into a bounded, credit-based push that behaves similarly to pull once tuned correctly. - Pure pull's own cost is **latency and polling overhead**: because the consumer only discovers new messages on its next poll call, there's an inherent delay between a message being produced and a consumer noticing it, bounded by how frequently it polls; poll too rarely and you add latency, poll too aggressively when there's nothing new and you waste network round-trips and add needless load on the broker for empty responses. - Many pull-based systems mitigate this with **long-polling** — the broker holds the connection open and only responds once data is available or a timeout elapses — which gets close to push-like low latency while keeping the consumer-initiated, rate-controlling shape of pull. ## Failure modes Failure modes differ correspondingly. - **A push system without adequate flow control** shows up as a consumer's process suddenly getting slammed — its receive buffer or in-memory queue balloons, garbage collection pressure spikes, and in the worst case the consumer process crashes from memory exhaustion or the broker's connection to it drops, at which point in-flight unacknowledged messages typically get redelivered to someone else — potentially worsening a cascading overload if the whole fleet is under load simultaneously. - **A pull system's failure mode is subtler**: a consumer that's polling too infrequently (misconfigured poll interval, or blocked on a slow downstream call between polls) looks identical, from a lag perspective, to a consumer that's just slow at processing — you see lag grow either way, and diagnosing whether the bottleneck is "processing time per message" versus "time between poll calls" requires separate metrics (processing latency vs. poll interval/fetch size) rather than lag alone. ## A concrete comparison A concrete comparison: Kafka's pull model lets a single broker serve thousands of consumers each polling at whatever rate matches their own processing capacity, without the broker needing to track per-consumer send-rate state — that scalability characteristic is a major reason Kafka favors pull for its high-throughput, log-based use case. By contrast, RabbitMQ's default push behavior, tuned with a `prefetch_count` of, say, 10, means the broker will push up to 10 unacknowledged messages to a given consumer and then stop pushing more until acks bring that count back down — effectively a credit-based hybrid that gets RabbitMQ's low per-message latency advantage from push while still bounding how far a slow consumer can fall behind on in-flight work before push traffic to it pauses.

  • How does a credit-based or prefetch mechanism turn a push system into something closer to pull-style back-pressure?
    It caps the number of unacknowledged (in-flight) messages the broker will send a given consumer before waiting for acks — once that cap is hit, the broker stops pushing more, effectively pausing until the consumer signals capacity by acking. That converts unbounded push into bounded, credit-gated push, which behaves much like pull's implicit rate control, just implemented as an explicit counter instead of a request-driven fetch.
  • Why might a system choose push despite its flow-control complexity?
    Push gives lower latency for low-to-moderate throughput, event-driven use cases where you want near-instant delivery the moment a message is ready, without paying the polling-interval delay pull introduces — webhooks and many task-queue systems favor this for responsiveness over raw throughput scalability.
  • Does long-polling fully eliminate the latency downside of pull?
    It substantially reduces it — the broker holds the request open and responds as soon as data arrives rather than making the consumer wait for its next scheduled poll — but it doesn't make it identical to push, since the consumer still must have an outstanding poll request active, and there's still per-request overhead compared to a broker proactively opening a send whenever it likes.

Push is like a waiter who keeps bringing you courses whenever the kitchen finishes them, whether or not you've finished the last one — great for speed, bad if you're a slow eater and plates pile up. Pull is like a buffet where you go back for more only when your plate is empty — you naturally never take more than you're ready for, but you might have to get up and check more often than you'd like if you eat fast.

saying these in an interview costs you the question

  • Claims push systems never need any flow control
  • Thinks pull inherently has higher throughput ceiling with no latency cost
  • Doesn't know what a prefetch/QoS credit limit does
  • Can't explain why pull back-pressure is 'implicit'

context