Why does streaming telemetry from network devices beat five-minute SNMP polling when you monitor 2,000 routers and switches?
answer
- who starts each exchange
- a five-minute average flattens bursts
- polls fail when the network is stressed
- subscribe once, the device keeps sending
- sampled on a timer, or on change
basics
~20 sSNMP polling makes a manager ask every device for every value each cycle, so data arrives late, averaged and lost under stress. Streaming telemetry subscribes once; each device then pushes YANG-modelled values on a timer or on change.
solid answer
~50 sWith polling, the manager drives everything: each cycle it sends requests (a `GetBulkRequest` to walk a table) to all 2,000 devices, and a five-minute cycle turns every counter into a 300-second average. RFC 8641 lists the costs: latency, cycles missed or delayed *when the network is under stress*, uneven intervals that are hard to compare, and wasted load for data that rarely changes. Streaming telemetry replaces the request loop with a **subscription**: the collector asks once, through gNMI `Subscribe` (an OpenConfig specification, not an RFC) or YANG-Push (RFC 8641), and the device sends values on its own schedule, each stamped with the time it was collected. Counters are sampled every few seconds; state such as `oper-status` is pushed the moment it changes. It is not free: the device and collector carry the load, and a publisher should decline a subscription it cannot keep.
go deeper
Recall who starts the exchange: in polling the manager asks every cycle; in streaming the collector subscribes once and the device keeps sending, on a timer or when a value changes.
Explain why a five-minute counter is a 300-second average, compute what a short burst looks like after averaging, and say which data suits sampled versus on-change delivery.
Show the operational argument: polls fail under exactly the stress you are diagnosing, device timestamps beat arrival times, and a publisher should refuse a subscription it cannot sustain.
Frame it as a cost move rather than a free win: decide which signals justify high-rate streaming, where SNMP stays good enough, and what the collector tier must absorb.
## The setting An operations team monitors **2,000 routers and switches** with SNMP. A central manager polls every device every five minutes for interface counters and status. Interviewers use exactly this picture to ask why the industry moved to **model-driven streaming telemetry**: devices pushing data described by YANG models instead of waiting to be asked. Two specifications carry the push model on network devices: - **gNMI** (gRPC Network Management Interface) - an **OpenConfig specification**, version 0.10.0, not an IETF RFC. Its `Subscribe` RPC lets a collector request a stream of values. - **YANG-Push** - the IETF's mechanism, **RFC 8641**, built on the subscription framework of **RFC 8639**, carried over NETCONF (RFC 8640) or RESTCONF (RFC 8650). ## What polling costs at scale In SNMP polling the **manager drives every exchange**. Each cycle it sends a request per device, often several: walking an interface table takes repeated `GetBulkRequest` PDUs (RFC 3416) because the manager cannot know in advance how many rows exist. RFC 8641's introduction names the problems directly: 1. **Latency** - a value is only as fresh as the last completed poll. 2. **Missed or delayed cycles** - requests get lost or late "often when the network is under stress and the need for the data is the greatest". 3. **Uneven intervals** - polls drift, so two samples are not exactly five minutes apart and rates are hard to calibrate. 4. **Wasted load** - asking 2,000 devices for status that changed nowhere is work for the network, the devices and the manager. The second point is the one that hurts during an incident: the moment links are congested or a device's CPU is busy is the moment polls time out. ## The averaging trap, worked through SNMP interface counters are cumulative: the manager subtracts two readings and divides by the time between them. That makes a five-minute poll a **300-second average**. Suppose a link runs at 10% utilisation except for a 20-second burst at 100%: | Interval | Seconds | Utilisation | |---|---|---| | Normal traffic | 280 | 10% | | Burst | 20 | 100% | | **Poll reports** | 300 | (280 x 10 + 20 x 100) / 300 = **16%** | The burst's octets are all counted, but they are diluted into a number that looks healthy. With a 10-second sample the same burst shows as two samples near 100%. The data was never missing; the resolution was. ## What a subscription changes A **subscription** is a contract: the collector states once which paths it wants and under which trigger, and the device keeps sending without further requests. RFC 8639 defines it as information receivers "wish to have pushed from the publisher without the need for further solicitation". - **Sampled or periodic** updates send current values on a fixed interval - gNMI `SAMPLE` with a `sample_interval`, YANG-Push `periodic` with a `period`. This suits counters. - **On-change** updates send a value only when it changes - gNMI `ON_CHANGE`, YANG-Push `on-change`. This suits state such as interface status, where a flap is reported in seconds rather than at the next poll. - **Timestamps come from the device.** A gNMI `Notification` carries the time the value was collected, in nanoseconds since the Unix epoch, so the collector no longer infers timing from when a response happened to arrive. - **Data is modelled.** Paths follow YANG models, so the same path means the same thing across devices that implement the model. ## What streaming does not fix Push moves cost; it does not erase it. - A one-second sample over every counter on every device can cost more than a five-minute poll. RFC 8641 says a publisher should reject a subscription it is unlikely to fulfil, and notes that one-second periods consume more resources than one-hour periods. - The collector must ingest a continuous stream from 2,000 devices and store far more points. - SNMP still has a role: SNMP traps and informs report defined events, and many tools still read MIBs. The usual pattern is to stream the high-rate data and keep SNMP where it is cheap and sufficient. The interview answer, in one line: polling makes the manager pay for every value every cycle and fails when you need it most; a subscription lets the device report at the right resolution, with its own timestamps, and report changes as they happen.
- SNMP already has traps. Why isn't that enough to replace polling?A trap reports one defined event, such as `linkDown`, with a few variable bindings, and an `SNMPv2-Trap-PDU` is not acknowledged, so a lost one is simply gone (an `InformRequest-PDU` is acknowledged). Traps cover events, not the continuous counter series you need for utilisation. A streaming subscription covers both: counters on a sampling interval and any modelled leaf on change.
- Is streaming always cheaper for the device than polling?No. Cost follows what you subscribe to and how often. A one-second sample of full interface tables is heavier than a five-minute poll. RFC 8641 says a publisher should reject a subscription it is unlikely to keep and notes that short periods consume more resources, so you size intervals per data type rather than streaming everything as fast as possible.
saying these in an interview costs you the question
- Streaming telemetry is just SNMP polling with a shorter interval.
- gNMI is an IETF RFC, the standards-track successor of SNMP.
- A five-minute counter average still shows every short traffic burst.
- Once devices push data, the collector no longer has a scaling problem.
- A device streams everything it knows without any subscription being set up.