How would you design SNMP notification delivery for thousands of devices so that a collector outage cannot silently hide a core link-down?
answer
- assume notifications get lost
- detect, then repair
- two receivers, separate failure domains
- ifLastChange tells what you missed
basics
~20 sAssume notifications will be lost and design to detect and repair: informs for critical events, two receivers in separate failure domains, receivers that deduplicate and store before acknowledging, and a reconciliation poll that catches what was missed.
solid answer
~40 sNo SNMP notification is guaranteed, so the goal is that loss is never silent. Classify events: informs only for the few whose loss matters, such as a core `linkDown`, traps for the rest. Send critical notifications to two receivers in different failure domains — RFC 3413's notification tables can select several targets, and each gets its own copy — and deduplicate downstream. Make receivers persist before they acknowledge. Bound the inform retry window to the outages you plan for, mindful of agent memory. Then reconcile: after any receiver gap and on a regular cycle, poll `sysUpTime` and `ifLastChange`; a changed `ifLastChange` reveals a transition even if the link is back up. Contain storms with agent rate-limiting and notification filtering. Finally, measure the pipeline: count transitions reconciliation finds that no notification reported.
go deeper
Recall that SNMP notifications can be lost, so a monitoring system also checks device state directly.
Explain how informs, multiple notification targets configured per RFC 3413 and a reconciliation poll each reduce the chance of missing an event.
Design receivers that store before acknowledging and deduplicate, and use ifLastChange and sysUpTime to find transitions no notification reported.
Own the trade-offs: device state versus receiver redundancy, reconciliation interval versus polling load, recovery bursts, and a measured loss rate as the pipeline's own health signal.
## Start from the failure model SNMP notifications travel as single UDP datagrams. A **trap** is never acknowledged; an **inform** is acknowledged and retried, but RFC 3416 still says there is "no guarantee of delivery". Collectors restart, links congest, agents rate-limit. At thousands of devices some notifications **will** be lost, so the design goal is not "never lose one" but **"never lose one without noticing, and recover the state anyway"**. ## The layers and what each buys | Layer | Closes | Costs | |---|---|---| | Informs for critical events | short receiver gaps; agent knows of failure | agent state per pending inform, retransmissions, recovery bursts | | Two receivers | one receiver down or isolated | double traffic, duplicates, double inform state | | Store-before-acknowledge receivers | crash between reply and storage | receiver write latency on the reply path | | Reconciliation poll | anything the layers above missed | polling load, detection delay of up to one cycle | | Rate limits and filtering | storms that overwhelm receivers | detail lost at the source | No single layer is enough; together they turn silent loss into bounded, measured delay. ## Delivery: what the device sends - **Classify notifications.** A core `linkDown`, a `coldStart` or a routing adjacency loss deserves an inform; a busy access port's flapping does not. RFC 3413's SNMP-NOTIFICATION-MIB lets one notification go as an inform to some targets and a trap to others (`snmpNotifyType` per entry, selecting SNMP-TARGET-MIB rows by tag). - **Two receivers.** A notification originator sends to **each** selected target (RFC 3413), so two receivers in different failure domains each get their own copy, and with informs each copy is retried on its own. This is the redundancy that covers a long outage of one collector, which no retry count can. - **Size the retries.** With a fixed wait `T` and retry count `R`, an inform is held for `(R + 1) × T`. Fit that to planned receiver outages, multiplied by the burst rate, against what the device can hold. - **Filter and rate-limit at the source.** RFC 2863 says an agent should limit the rate of link traps when an interface flaps, and RFC 3413's notification filtering lets a target receive only the notifications it needs. ## Receivers: what the collector does 1. **Persist before acknowledging.** RFC 3416 has the receiver present the inform to its application and then respond; make the application write durably first, so an acknowledgement means stored. 2. **Deduplicate.** Two receivers and inform retransmissions both produce copies. Key on the sender plus `sysUpTime.0`, `snmpTrapOID.0` and the event's objects, and make alerting idempotent. 3. **Stamp arrival time.** `sysUpTime.0` is uptime in hundredths of a second, not wall time. ## Reconciliation: repair what was missed The IF-MIB gives the poll a precise signal. **`ifLastChange`** is the value of `sysUpTime` when the interface entered its current operational state, so: - if `ifLastChange` differs from the last value recorded, a transition happened — **even if** `ifOperStatus` reads `up` again now; - if `sysUpTime` is lower than last time, the agent re-initialised, and `ifLastChange` values from before that read zero, so resynchronise the device fully. Poll after every known receiver gap and on a steady cycle. The cycle is the trade-off: 5,000 devices × (1 `sysUpTime` + 48 `ifLastChange`) = 245,000 values; every 5 minutes that is about **817 values per second** across the estate, and a missed transition is found within 5 minutes at worst. ## Measure the pipeline itself - **Loss rate**: transitions found by reconciliation with no matching notification. A rising count is the pipeline failing before any outage proves it. - **Inform outcomes**: agents know which informs failed; export or poll that where the device exposes it. - **Receiver liveness**: a receiver that has heard nothing from a busy estate for longer than usual is itself suspect. Some operators add a periodic synthetic notification as a heartbeat — an operating practice, not an SNMP mechanism. ## Trade-offs a lead owns - **Informs versus receivers.** Informs push cost onto every device; a second receiver pushes it onto infrastructure you control. In a large estate of constrained devices, two receivers plus reconciliation often beat long inform retries. - **Reconciliation interval versus load.** A shorter cycle finds loss sooner and costs polling capacity; critical devices can be polled more often than the rest. - **Recovery bursts.** When a collector returns, retrying informs and the reconciliation sweep arrive together; receivers must be sized for that moment. ## The interview point A strong answer refuses the premise that the right PDU makes notifications reliable, layers detection and repair, and puts numbers on the trade-offs instead of naming tools.
- Why does polling ifOperStatus alone miss a link-down that happened during a collector outage?Because the link may have come back before the poll. ifOperStatus reports only the current state, so a down-then-up between two polls reads up both times. ifLastChange records the sysUpTime at which the interface entered its current state; if it moved since the last poll, a transition happened in between, which is exactly the event the lost notification would have reported.
- When would you choose two trap receivers over informs to one receiver?When the devices are many and constrained, and the outage you fear is a receiver failure rather than a lost packet. Informs cost every device state and retransmissions and still fail if the receiver is down longer than the retry window. Two trap receivers in different failure domains cover one being down at any length, cost the devices only an extra datagram, and reconciliation catches what both missed.
- How do you know the notification pipeline is losing events before an outage proves it?Compare the two independent views. Every transition the reconciliation poll finds through a changed ifLastChange should have a matching notification at a receiver; one without a match is a measured loss. Track that count per device and per receiver over time, and alert on it, so a collector that silently stopped receiving, or an agent that stopped sending, shows up as a rising gap.
saying these in an interview costs you the question
- Switching every notification to informs makes the pipeline reliable.
- Two receivers mean each event is delivered exactly once.
- Polling ifOperStatus after an outage reveals every missed transition.
- Retry counts can be raised until no outage loses an inform.
- If no notifications arrive, the network must be healthy.