What does push-based metrics collection buy you, and what does the backend lose by never asking?
answer
- which side opens the connection
- no inbound route, no discovery needed
- the backend only knows what arrived
- silence has half a dozen causes
- you must supply the expected-sender list yourself
basics
~20 sPush needs no inbound route and no discovery - a sender exists the moment it sends, which is what makes it work from behind address translation and from very short-lived processes. The cost: silence is ambiguous, so died and idle look identical.
solid answer
~50 sIn a push model the instrumented process, or an agent beside it, sends measurements to a fixed ingest endpoint. That buys three things: **no inbound network path is needed**, so a sender behind address translation works with egress-only rules; **no discovery is needed**, because a new instance exists the moment it sends; and **very short-lived processes are covered**, since nothing has to be alive when a collector calls. What the backend gives up is any ability to reason about absence. It knows only what arrived, so a gap could be a crash, a lost network, a wrong endpoint in a config, a deployment that never happened, or a genuinely idle service. Identity is self-declared too. Getting liveness back means keeping an independent list of who should report and alerting on how long each has been quiet.
code
pseudocode · 9 linesqueue = bounded_queue(capacity=10000)
every flush_interval:
batch = queue.drain(max=500)
if not send(ingest_endpoint, batch):
retry_or_drop(batch) # this choice IS your delivery guarantee
on_shutdown:
flush(queue) # never runs if the process is killed outrightgo deeper
Recall the direction of travel: in a push model the application sends to a fixed endpoint, so only an outbound network path is needed. Know the headline trade - easy to deliver from anywhere, but the receiver cannot tell a dead sender from a quiet one.
Explain the mechanics: no discovery, egress-only connectivity, self-declared identity, and per-sender interval configuration. Be able to list several distinct causes of a gap in pushed data and say why the backend cannot separate them.
Demonstrate the repair. Describe sourcing an expected-sender roster from deployment records, alerting on last-seen age rather than absence, recording what the ingest tier itself shed, and defending against a sender that suddenly multiplies its dimensions.
Own the framing that push trades a network requirement for a knowledge requirement. Be ready to say who in the organisation maintains the roster of expected senders, and what it means for incident response when that list drifts out of date.
## What changes when the sender initiates In a push model the direction of the connection flips: the instrumented process - or a shipper running beside it - opens a connection outward and sends its measurements to an ingest endpoint. The backend is a receiver. It never asks, never holds a list of who is expected, and learns of a sender's existence from the first payload that arrives. StatsD-style line protocols, OTLP and a metrics agent writing into a hosted backend are all this same shape. ## What push buys - **No inbound reachability requirement.** Only an outbound path is needed, so senders behind address translation, inside a customer's network, on a mobile or edge device, or in an environment where nobody will open a hole for a monitoring system all work. Egress-only firewall rules are far easier to get approved than inbound ones. - **No discovery requirement.** A new instance does not have to be found and added to any list; it participates by sending. For a fleet that scales in and out constantly, that removes an entire class of "the new pods are not being collected" incident. - **Ephemeral workloads are covered.** A process that lives for a few seconds can still report, because reporting does not depend on being alive at some external moment of collection. - **The sender paces itself.** It can buffer, batch and compress on its own terms, which suits a device on an expensive or intermittent link far better than answering a poll on someone else's schedule. - **Boundaries are crossable in one direction.** Reporting out of an environment you do not control is often the only option available at all. ## The bill: three things the backend gives up **1. Absence becomes ambiguous.** This is the central cost and it is not fixable inside the model. A receiver knows what arrived; it cannot know what should have arrived. | A gap in pushed data could mean | How the backend distinguishes it | | --- | --- | | The process crashed | It cannot | | The host lost its network path | It cannot | | The endpoint in the config is wrong | It cannot | | The service was never deployed at all | It cannot | | The service is running and simply idle | It cannot | | The ingest tier shed the data under load | Only from its own side, if it records that | All six look the same from the receiving end: nothing. That is why alerting on "no data" in a push-only system is either noisy or silent - noisy if idleness is normal, silent if the sender was never there to be missed. **2. Identity is self-declared.** The sender chooses the name and the dimensions its data is filed under. A copy-pasted configuration can make a staging instance report as production; a bad release can put a per-request value into an identifying dimension and create a flood of new series with no chokepoint upstream to stop it. The backend needs ingest-side limits and quotas precisely because it cannot trust what it is told. **3. Interval control is distributed.** Resolution now lives in every sender's configuration. Changing it is a fleet-wide rollout, mismatched settings mean different services are measured at different granularity, and a synchronized restart of many senders can arrive as one correlated spike on the ingest tier. ## Getting the liveness signal back The standard repair is to reintroduce, by hand, the thing a polling collector had for free: 1. Keep an **independent roster** of who is expected to report, sourced from deployment records, the orchestrator or a service catalogue - not from the telemetry itself, which is circular. 2. Derive a **last-seen timestamp** per expected sender and alert on its **age**, not on the absence of a series. Age is a real signal even when a service is legitimately idle, because a healthy sender still reports something. 3. Give every sender a cheap always-emitted heartbeat so that "quiet" and "gone" stop looking alike. 4. Record what the **ingest tier itself** rejected or shed, so that data lost on the receiving side is not misread as a sender problem. Notice what this list is: an inventory, a staleness signal and a health record - the exact three things the polling model hands over without being asked. The useful conclusion for an interview is that push and pull are not better and worse; they trade a network requirement for a knowledge requirement. Push removes the need to reach every process, and in exchange you take on the job of knowing what should be reporting.
- How would you make a missing sender alertable in a push-only setup?Keep a roster of who is expected to report, taken from deployment or orchestrator records rather than from the telemetry itself, join it against each sender's last-seen timestamp, and alert when that timestamp gets too old. It works, but note what you have done: rebuilt the inventory and the liveness check a polling collector provides for free, and made your alerting only as accurate as the roster.
- Why is self-declared identity a cardinality risk on the receiving side?Because the sender picks its own name and dimensions, one bad release that puts a request id, a customer id or a full URL into an identifying dimension starts creating new series at whatever rate that value changes. There is no target list acting as a chokepoint, so the first thing that notices is the backend. Ingest-side limits, per-tenant quotas and dimension allowlists are the defence.
- Does push guarantee lower latency between an event and its appearance on a dashboard?Not by itself. A sender that batches on a fifteen-second flush and a collector that polls every fifteen seconds have comparable end-to-end delay. Push can be made faster by flushing more often, but that trades away batching efficiency, and the sender's queue, its retry policy and the ingest tier's own buffering all add delay that a shorter flush does not remove.
A push model is a postcard from the field: plenty arrive, but a postcard nobody ever wrote looks exactly like one that was lost in the post.
saying these in an interview costs you the question
- Calls push real-time and pull delayed, with nothing behind the claim
- Alerts on absence of pushed data with no roster of expected senders
- Assumes a killed process still flushes its buffered measurements
- Thinks push removes the need to know what should be running
- Believes the backend can verify who a sender claims to be
- Treats a quiet service and a dead service as the same observation