skip to content

How do you scale pull-based and push-based metrics collection in one estate, and where is the line between them?

level: principalimportance: should knowfreq 46%

answer

  1. the two sides run out of room differently
  2. interval against targets times work per poll
  3. bounded queue, stated drop policy
  4. sharding needs a query layer that merges
  5. reachability and lifetime decide the line

basics

~20 s

Polling scales by budgeting the collection interval against target count and work per target, then sharding the collection tier per network domain. Sending scales by batching, bounded queues and an explicit drop policy. Draw the line by reachability and lifetime.

solid answer

~50 s

The two sides run out of room for different reasons. A polling tier's budget is **interval against targets times work per target**: the per-target timeout must fit inside the interval, and when the slowest targets approach that budget you shard the target set - which means owning an assignment and adding a query layer that merges the shards. It also has a topology constraint, because a collector must route to everything it polls, so you end up with one tier per network domain. A sending tier is limited by **queues and delivery**: batching amortises overhead but enlarges each loss, the in-process queue must be bounded with a stated drop policy, and retries turn a brief ingest failure into a correlated stampede without backoff and jitter. Poll what you can route to and that outlives a cycle; send from everything else.

go deeper

for a junior

Recall that neither model is free at scale: polling costs work on every cycle for every target, and sending costs memory in each sender's queue. Know that many estates run both rather than choosing one.

for a middle

Explain the levers on each side - interval, per-target timeout and target count when polling; batch size, flush interval, queue bound and retry policy when sending - and what the first symptom of overload looks like in each case.

for a senior

Show you have operated both. Talk about a timeout that no longer fits the interval, a retry stampede after an ingest outage, an unbounded queue that took down the application it was measuring, and what sharding a collection tier did to your queries.

for a principal

Own the boundary and its consequences: which workloads may send, who maintains the roster of expected senders, what identity conventions both intakes must obey, and how the estate proves months later that a control was actually running.

## Two different scaling budgets It is tempting to argue that one model "scales better". They do not fail at the same thing, and a lead is expected to name what limits each. | Dimension | Polling tier | Sending tier | | --- | --- | --- | | What limits it | Interval against targets times work per poll | Ingest throughput and per-sender queue memory | | What you tune | Interval, per-target timeout, target assignment | Batch size, flush interval, queue bound, retry policy | | What fails first | Collections overrun the interval | Queues fill and data is dropped at the sender | | What the failure looks like | Gaps in otherwise healthy series | Silence that the backend cannot explain | | Where the cost lands | On the targets, on every cycle | On the ingest tier, in correlated bursts | ## Scaling the polling side The arithmetic is unforgiving but simple. In a 41-service estate averaging 2,847 series per instance, the work per cycle is the number of targets multiplied by the cost of rendering and parsing that payload, and all of it must complete comfortably inside the interval. Three consequences follow. 1. **The per-target timeout must fit inside the interval.** If it does not, collections overlap, resolution silently degrades, and the first symptom is holes in the data for the slowest services rather than an explicit error. 2. **Growth is handled by sharding the target set**, and sharding is where the difficulty is. Assigning targets across collectors is a coordination problem, and once data lives on more than one collector every query needs a layer that fans out and merges - which is a whole component you did not have when one process held everything. 3. **Network topology puts a floor on the number of tiers.** A collector must open connections to everything it polls, so each network domain that cannot be entered from outside needs its own collection tier. That is an operational cost that has nothing to do with data volume. Remember also that the cost falls partly on the targets: every instance spends CPU rendering its current values on every cycle, whether anyone looks at them or not. ## Scaling the sending side Here the questions are queueing questions, and answering them vaguely is the usual failure. - **Batching** amortises connection and framing overhead, at the price of delay and of blast radius: a lost batch loses everything in it, so larger batches mean coarser loss. - **The queue must be bounded, with a stated policy for what happens when it fills** - drop the oldest, drop the newest, or apply back pressure to the application. Unbounded queueing is not reliability; it is an out-of-memory failure moved to a worse moment, and blocking the application means telemetry can take down the service it was meant to observe. - **The delivery guarantee is a choice.** Fire-and-forget loses data quietly; retrying produces duplicates the backend must tolerate. Both are defensible, but the choice must be explicit, and retries need backoff and jitter or a two-minute ingest blip becomes a synchronised stampede from every sender at once. - **Size the ingest tier for the correlated peak.** Senders are not independent: a fleet-wide deployment restarts everything together and a recovering backend gets everyone's backlog simultaneously. Per-tenant quotas and load shedding are what stop one misbehaving service from denying the rest. ## Where the line falls in practice Two properties decide it, and neither is a matter of taste: 1. **Can the collector route to it?** If not - a partner-hosted scheduling gateway, an agent on a court clerk's device, anything behind a boundary nobody will open - it sends. 2. **Does it outlive a collection cycle?** If not - nightly reconciliation, import jobs, anything that finishes in seconds - it sends. Everything else is polled, because the inventory and per-cycle liveness signals are worth more than the convenience of sending. In a 41-service estate that typically lands at roughly seven senders and the rest polled inside two collection tiers, with both intakes normalised into one store. What must be identical across the line: **identity conventions** (service, environment and instance dimensions spelled the same way on both sides, or the halves cannot be queried as one estate), unit and naming conventions, and retention. A dashboard built on one convention will silently cover only half the estate, and nobody notices until the half it omits is the half that broke. ## The part people forget: evidence The two intakes are not equally accountable, and that matters when someone asks for six months of proof rather than a live dashboard. A polling tier records its own failures, so a gap in the data comes with an explanation of why the data is missing. A sending tier's gap is unexplained by construction, and after the fact it cannot be distinguished from a period when nothing happened. If an estate has to demonstrate that a control was running continuously, the pushed half needs a deliberately maintained roster of expected senders plus staleness records kept for the whole retention window - otherwise the honest answer to "prove it was running in March" is that you cannot.

  • Which single number tells you a polling tier is running out of headroom?
    The ratio of the time a collection takes to the interval it must fit in, tracked per target and at the worst case rather than the average. While the slowest targets sit at a small fraction of the interval there is room; as they approach it you are one deployment away from gaps. It leads the collector's own CPU and memory as an indicator, because the arithmetic breaks before the process does.
  • Why is sharding a polling tier harder than adding another ingest node?
    An ingest node is largely interchangeable: put another behind the load balancer and senders are unaffected. A polling tier owns an assignment of targets, so adding a collector means redistributing that assignment consistently and without gaps or double collection. Afterwards the data lives in more than one place, so queries need a component that fans out and merges results - a piece of architecture the ingest side never needed.
  • If both intakes feed one store, what has to be identical on both sides?
    The identity dimensions above all - service, environment and instance named and spelled the same way - because that is what lets a query span both halves. Then units and metric naming, so the same quantity is not recorded two ways, and retention, so a comparison across the line does not silently end early on one side. Divergence here is usually discovered during an incident, at the worst moment.

saying these in an interview costs you the question

  • Claims one model simply scales better without naming what limits either
  • Adds collectors without deciding how targets are assigned
  • Retries pushed batches indefinitely with no backoff or jitter
  • Runs an unbounded in-process queue and calls it reliability
  • Assumes one collection tier can route into every network in the estate
  • Lets the two intakes use different identity conventions in one store