skip to content

Push vs Pull Collection

The two ways metrics reach a backend and why the choice matters for short-lived jobs, NAT, and knowing a target is down. A standard architecture question when comparing Prometheus-style scraping to push agents.

on this pageshow

questions

4

What does a pull-based metrics collector get for free, and what does it demand of every target?

level: middleimportance: must knowfreq 74%

answer

  1. who starts the conversation
  2. the collector already has a list
  3. a failed attempt is data too
  4. one config sets everyone's resolution
  5. targets must be routable and always answerable

basics

~20 s

A pull-based collector already holds the list of what should exist, so every failed poll is itself a liveness signal, and one setting fixes the whole fleet's sampling resolution. In exchange each target must stay routable and answerable at any instant.

solid answer

~50 s

Pulling inverts who initiates: the collector holds a list of **targets** and polls each on a schedule, so three things arrive for free. First an **inventory** - the collector knows what ought to answer, so something that should be there and is not becomes an alertable condition instead of silence. Second a **liveness verdict every cycle** - a refused connection, a timeout or a bad payload is recorded per target, which is why a scraper can tell you a process is down with no heartbeat plumbing at all. Third **one place that fixes resolution**, because the interval lives in the collector's configuration rather than in dozens of independently deployed services. The targets pay for this: each must be routable inbound from the collector, must answer inside the per-target timeout, and must hold current values in memory so that any arbitrary moment is a valid read.

code

pseudocode · 9 lines
pseudocode
every collection_interval:
    for target in target_list:            # the list is the inventory
        started = now()
        try:
            payload = http_get(target.url, timeout=target_timeout)
            store(target.identity, parse(payload), at=started)
            store(target.identity, reachable=true, at=started)
        except timeout, refused, parse_error:
            store(target.identity, reachable=false, at=started)   # liveness, for free

go deeper

for a junior

Be ready to say who initiates in each model: in pull the monitoring system connects to the application on a schedule; in push the application sends. Know that a pull system is configured with a list of things to poll.

for a middle

Explain the mechanics behind the free signals: the target list is an inventory, every poll attempt yields a recorded success or failure, and the interval lives in one configuration. Then name the price - inbound routability, an endpoint answerable at any instant, and a response inside the timeout.

for a senior

Show the operational judgment: what a per-target timeout that approaches the interval does to resolution, what rendering thousands of series costs a busy process on every cycle, and how you treat a metrics endpoint as network surface that needs authorization.

for a principal

Own the estate-wide argument. Be able to say why the inventory and liveness signals are the reason polling stays the default for long-lived services, and to state honestly which classes of workload can never meet those demands and therefore need a different intake.

## Who initiates decides what you know In a pull model the **collector** initiates. It holds a set of **targets** - endpoints it was configured with or discovered - and on a fixed cycle it opens a connection to each one, asks for the current values, and stores what comes back against the time of the request. The instrumented process is passive: it keeps its counters and gauges in memory and renders them on demand, conventionally over HTTP at a path such as `/metrics`. That inversion is the whole subject. Because the collector does the asking, it must already know **who to ask**, and that list - which looks at first like pure configuration overhead - is the source of everything the model gives you for nothing. ## Three by-products you would otherwise have to build 1. **An explicit inventory of what should exist.** The target list is a statement of intent: these processes are supposed to be answering. That makes *absence* a first-class, alertable condition rather than a gap nobody notices. In a model where senders announce themselves, a service that never starts and a service that was never deployed look identical - both are simply not there. 2. **A liveness verdict on every cycle.** Every poll has an outcome: it connected or it did not, it answered inside the timeout or it did not, the payload parsed or it did not. A pull collector records that outcome per target, per cycle - Prometheus, for example, writes a synthetic `up` sample carrying 1 for a successful poll and 0 for a failed one, right beside whatever the target returned. You get per-process reachability without a heartbeat, without a health-check pipeline and without any extra moving part. 3. **One place that controls the sampling interval.** Resolution is a property of the collector, not of the code being measured. Halving the interval for one environment is a change in one configuration, not a rollout across every service that reports; and because the collector paces its own work, it can spread load rather than being hit by whatever cadence dozens of teams happened to pick. A fourth, quieter one: **identity is assigned at the collection side**. What the data is called comes from the target record the collector already holds, not from a string the process chose for itself, so a misconfigured deployment cannot quietly file its numbers under another service's name. | Free by-product | What you build in its absence | | --- | --- | | Target inventory | A registry every process must join, kept honest by hand | | Per-cycle liveness | A heartbeat signal plus staleness alerting for every sender | | Central interval control | A fleet-wide config rollout every time resolution changes | | Collection-side identity | Trust that each sender names and labels itself correctly | ## What every target owes in return - **Routability, inbound.** The collector opens the connection, so a network path must exist *into* the target's environment. Address translation, per-namespace firewall policy, a partner's network or a client device all break this, and a target the collector cannot reach is invisible no matter how well instrumented it is. - **Answerability at any instant.** The target does not choose when it is read, so every moment must be a valid read: values held ready in memory, nothing that only makes sense as "what happened since you last asked", and no expensive computation on the request path. - **Survival across at least one cycle.** A process that starts and exits between two polls is never observed. It is not that its data is delayed - it never existed as far as the collector is concerned. - **A response inside the timeout.** The per-target timeout has to fit inside the interval, otherwise collections overlap and the effective resolution silently degrades. A target that takes longer than its budget produces gaps, not slower data. - **The cost of being read.** Rendering the current value of every series costs the target CPU and allocation on every cycle. At a few thousand series per instance across a few dozen services this is small but real, and it grows with the series count rather than with traffic. - **A defended endpoint.** An always-answerable endpoint is network surface. It needs authorization or network policy, because what it exposes - internal names, versions, queue depths, error counts - is a useful map for anyone who reaches it. ## Where the demands bite The list above is, almost exactly, the list of situations where teams reach for push instead: processes too short-lived to be caught, workloads behind a boundary the collector cannot cross, clients and edge devices with no inbound route at all, and third-party systems nobody will let you poll. Pull remains the default for long-lived server-side services precisely because those services can meet its demands cheaply - and because the inventory and liveness signals it hands over free are the two things a receive-only backend can never reconstruct on its own.

  • If the collector already reports whether a target answered, why do teams still run a separate external check?
    Because the poll only proves reachability *from where the collector stands*, which is usually inside the same network. It says nothing about the path a real user takes - DNS, the public load balancer, TLS at the edge - and it cannot report on itself: a collector that is down produces the same silence as a healthy estate. An outside-in check answers a different question, and the two are not substitutes.
  • Which timestamp does a pulled value carry, and why does that matter?
    The collector timestamps at collection time, so a value's age is bounded by the interval and every target on the same cycle lines up on a comparable grid. The risk is a target that computes values lazily when asked: the numbers then describe a moment slightly before the timestamp, and an expensive computation on the request path widens that skew exactly when the system is busiest.
  • What happens when a target's response consistently takes longer than its collection timeout?
    You get gaps, not slow data. The attempt is abandoned and recorded as a failure, so the series has holes while the process is in fact alive and serving traffic - a state that reads as an outage on a dashboard. The fixes are to cut what the endpoint renders, make rendering cheaper, or give that target a longer budget inside a longer interval.

It is a roll call rather than a suggestion box: because the register is written in advance, a name that answers nothing is itself the finding.

saying these in an interview costs you the question

  • Claims pull is universally more scalable than push, with no numbers
  • Thinks the collector discovers new instances with no target list or discovery at all
  • Cannot say what a collector records when a target times out
  • Believes each service decides how often it is sampled
  • Assumes polling needs no inbound network path into the target's environment
  • Treats an always-answerable metrics endpoint as costing the target nothing
open as a page

What does push-based metrics collection buy you, and what does the backend lose by never asking?

level: juniorimportance: should knowfreq 56%

basics

~20 s

Push needs no inbound route and no discovery - a sender exists the moment it sends, which is what makes it work from behind address translation and from very short-lived processes. The cost: silence is ambiguous, so died and idle look identical.

open as a page

How do you scale pull-based and push-based metrics collection in one estate, and where is the line between them?

level: principalimportance: should knowfreq 46%

basics

~20 s

Polling scales by budgeting the collection interval against target count and work per target, then sharding the collection tier per network domain. Sending scales by batching, bounded queues and an explicit drop policy. Draw the line by reachability and lifetime.

open as a page

How do you get metrics out of a batch job that exits before any collector could poll it?

level: seniorimportance: nice to knowfreq 42%

basics

~20 s

Three options: push to an intermediary that holds the last value for later collection, push straight to a receive-capable backend before exiting, or have a long-lived process report on the job's behalf. Each trades away either freshness or correct attribution.

open as a page