What does a pull-based metrics collector get for free, and what does it demand of every target?
answer
- who starts the conversation
- the collector already has a list
- a failed attempt is data too
- one config sets everyone's resolution
- targets must be routable and always answerable
basics
~20 sA pull-based collector already holds the list of what should exist, so every failed poll is itself a liveness signal, and one setting fixes the whole fleet's sampling resolution. In exchange each target must stay routable and answerable at any instant.
solid answer
~50 sPulling inverts who initiates: the collector holds a list of **targets** and polls each on a schedule, so three things arrive for free. First an **inventory** - the collector knows what ought to answer, so something that should be there and is not becomes an alertable condition instead of silence. Second a **liveness verdict every cycle** - a refused connection, a timeout or a bad payload is recorded per target, which is why a scraper can tell you a process is down with no heartbeat plumbing at all. Third **one place that fixes resolution**, because the interval lives in the collector's configuration rather than in dozens of independently deployed services. The targets pay for this: each must be routable inbound from the collector, must answer inside the per-target timeout, and must hold current values in memory so that any arbitrary moment is a valid read.
code
pseudocode · 9 linesevery collection_interval:
for target in target_list: # the list is the inventory
started = now()
try:
payload = http_get(target.url, timeout=target_timeout)
store(target.identity, parse(payload), at=started)
store(target.identity, reachable=true, at=started)
except timeout, refused, parse_error:
store(target.identity, reachable=false, at=started) # liveness, for freego deeper
Be ready to say who initiates in each model: in pull the monitoring system connects to the application on a schedule; in push the application sends. Know that a pull system is configured with a list of things to poll.
Explain the mechanics behind the free signals: the target list is an inventory, every poll attempt yields a recorded success or failure, and the interval lives in one configuration. Then name the price - inbound routability, an endpoint answerable at any instant, and a response inside the timeout.
Show the operational judgment: what a per-target timeout that approaches the interval does to resolution, what rendering thousands of series costs a busy process on every cycle, and how you treat a metrics endpoint as network surface that needs authorization.
Own the estate-wide argument. Be able to say why the inventory and liveness signals are the reason polling stays the default for long-lived services, and to state honestly which classes of workload can never meet those demands and therefore need a different intake.
## Who initiates decides what you know In a pull model the **collector** initiates. It holds a set of **targets** - endpoints it was configured with or discovered - and on a fixed cycle it opens a connection to each one, asks for the current values, and stores what comes back against the time of the request. The instrumented process is passive: it keeps its counters and gauges in memory and renders them on demand, conventionally over HTTP at a path such as `/metrics`. That inversion is the whole subject. Because the collector does the asking, it must already know **who to ask**, and that list - which looks at first like pure configuration overhead - is the source of everything the model gives you for nothing. ## Three by-products you would otherwise have to build 1. **An explicit inventory of what should exist.** The target list is a statement of intent: these processes are supposed to be answering. That makes *absence* a first-class, alertable condition rather than a gap nobody notices. In a model where senders announce themselves, a service that never starts and a service that was never deployed look identical - both are simply not there. 2. **A liveness verdict on every cycle.** Every poll has an outcome: it connected or it did not, it answered inside the timeout or it did not, the payload parsed or it did not. A pull collector records that outcome per target, per cycle - Prometheus, for example, writes a synthetic `up` sample carrying 1 for a successful poll and 0 for a failed one, right beside whatever the target returned. You get per-process reachability without a heartbeat, without a health-check pipeline and without any extra moving part. 3. **One place that controls the sampling interval.** Resolution is a property of the collector, not of the code being measured. Halving the interval for one environment is a change in one configuration, not a rollout across every service that reports; and because the collector paces its own work, it can spread load rather than being hit by whatever cadence dozens of teams happened to pick. A fourth, quieter one: **identity is assigned at the collection side**. What the data is called comes from the target record the collector already holds, not from a string the process chose for itself, so a misconfigured deployment cannot quietly file its numbers under another service's name. | Free by-product | What you build in its absence | | --- | --- | | Target inventory | A registry every process must join, kept honest by hand | | Per-cycle liveness | A heartbeat signal plus staleness alerting for every sender | | Central interval control | A fleet-wide config rollout every time resolution changes | | Collection-side identity | Trust that each sender names and labels itself correctly | ## What every target owes in return - **Routability, inbound.** The collector opens the connection, so a network path must exist *into* the target's environment. Address translation, per-namespace firewall policy, a partner's network or a client device all break this, and a target the collector cannot reach is invisible no matter how well instrumented it is. - **Answerability at any instant.** The target does not choose when it is read, so every moment must be a valid read: values held ready in memory, nothing that only makes sense as "what happened since you last asked", and no expensive computation on the request path. - **Survival across at least one cycle.** A process that starts and exits between two polls is never observed. It is not that its data is delayed - it never existed as far as the collector is concerned. - **A response inside the timeout.** The per-target timeout has to fit inside the interval, otherwise collections overlap and the effective resolution silently degrades. A target that takes longer than its budget produces gaps, not slower data. - **The cost of being read.** Rendering the current value of every series costs the target CPU and allocation on every cycle. At a few thousand series per instance across a few dozen services this is small but real, and it grows with the series count rather than with traffic. - **A defended endpoint.** An always-answerable endpoint is network surface. It needs authorization or network policy, because what it exposes - internal names, versions, queue depths, error counts - is a useful map for anyone who reaches it. ## Where the demands bite The list above is, almost exactly, the list of situations where teams reach for push instead: processes too short-lived to be caught, workloads behind a boundary the collector cannot cross, clients and edge devices with no inbound route at all, and third-party systems nobody will let you poll. Pull remains the default for long-lived server-side services precisely because those services can meet its demands cheaply - and because the inventory and liveness signals it hands over free are the two things a receive-only backend can never reconstruct on its own.
- If the collector already reports whether a target answered, why do teams still run a separate external check?Because the poll only proves reachability *from where the collector stands*, which is usually inside the same network. It says nothing about the path a real user takes - DNS, the public load balancer, TLS at the edge - and it cannot report on itself: a collector that is down produces the same silence as a healthy estate. An outside-in check answers a different question, and the two are not substitutes.
- Which timestamp does a pulled value carry, and why does that matter?The collector timestamps at collection time, so a value's age is bounded by the interval and every target on the same cycle lines up on a comparable grid. The risk is a target that computes values lazily when asked: the numbers then describe a moment slightly before the timestamp, and an expensive computation on the request path widens that skew exactly when the system is busiest.
- What happens when a target's response consistently takes longer than its collection timeout?You get gaps, not slow data. The attempt is abandoned and recorded as a failure, so the series has holes while the process is in fact alive and serving traffic - a state that reads as an outage on a dashboard. The fixes are to cut what the endpoint renders, make rendering cheaper, or give that target a longer budget inside a longer interval.
It is a roll call rather than a suggestion box: because the register is written in advance, a name that answers nothing is itself the finding.
saying these in an interview costs you the question
- Claims pull is universally more scalable than push, with no numbers
- Thinks the collector discovers new instances with no target list or discovery at all
- Cannot say what a collector records when a target times out
- Believes each service decides how often it is sampled
- Assumes polling needs no inbound network path into the target's environment
- Treats an always-answerable metrics endpoint as costing the target nothing