Monitoring 300 campus switches with SNMP, why poll on a five-minute loop and also take link-down notifications, and what does each one miss?
answer
- freshness versus completeness
- what happens between two polls
- a silent agent sends nothing
- the notification's path to the manager
- RFC 1157: events guide the polling
basics
~20 sSNMP polling gives complete, regular data and proves each agent still answers, but sees a failure up to one interval late and misses flaps between polls. Notifications arrive within seconds but only on events and can be lost, so estates use both.
solid answer
~50 sA five-minute SNMP poll of 300 switches reads every interface's counters and status on a schedule: it yields the deltas utilisation needs, and a timeout is itself evidence that a switch or its path is down. It costs load - with 52 interfaces and 4 objects each, 62,400 values per cycle, about 208 a second - and a failure is seen up to 300 seconds late, while a 20-second flap between polls leaves `ifOperStatus` reading `up` both times. A `linkDown` notification reaches UDP 162 within seconds, but it is sent only when the agent notices an event, a dead switch sends nothing, and a switch whose only uplink failed cannot deliver it at all. RFC 1157 describes the intended mix: monitoring "primarily by polling", with a limited number of traps guiding the timing and focus of the polling. A common practice is to let each notification trigger an immediate targeted poll of that switch.
go deeper
Recall that polling asks on a schedule and notifications arrive when something happens, and that most estates use both.
Explain what each flow carries - Get requests to UDP 161, linkDown to UDP 162 - and why counters for utilisation can only come from polling.
Quantify the trade-off for a real estate: values per second, worst-case detection delay, flaps hidden between polls, silent failures and notifications stranded behind the failed uplink.
Set the poll interval and notification policy as a budget across manager capacity, agent CPU and detection targets, and decide where independent paths for notifications are worth paying for.
## Two ways a manager learns about a switch An SNMP **manager** can learn the state of a switch's **agent** in two ways: - **Polling**: the manager sends `GetRequest`, `GetNextRequest` or `GetBulkRequest` messages to the agent's UDP 161 on a schedule and reads the answers. - **Notifications**: the agent's notification originator sends an unsolicited trap or inform to the manager's UDP 162 when an event occurs - for an interface, IF-MIB's `linkDown` and `linkUp` (RFC 2863), sent when an interface's `ifOperStatus` is about to enter the `down` state and when it has left it. The scenario: 300 campus switches, each with 48 access ports and 4 uplinks (52 interfaces), a manager polling every five minutes. ## What polling costs and what it misses Assume four objects per interface (`ifHCInOctets`, `ifHCOutOctets`, `ifInErrors`, `ifOperStatus`): 1. Per switch: 52 x 4 = **208** values per cycle. 2. For the estate: 300 x 208 = **62,400** values per cycle. 3. Per second over a 300-second cycle: 62,400 / 300 = **208** values a second. 4. If each response carries about 40 values, that is 62,400 / 40 = 1,560 request-response exchanges per cycle, about **5 a second**. That load is modest, and it buys a lot: every interface's counters at a known interval (the deltas any utilisation graph needs), and **proof of life** - a switch that stops answering is noticed at the next poll even if it never said anything. What polling misses: - **Latency.** A link that fails just after a poll is seen up to **300 seconds** later; on average about 150 seconds. - **Short flaps.** A port down for 20 seconds between two polls reads `up` on both. Among the status objects, only `ifLastChange` - the `sysUpTime` value when the interface entered its current state - records the flap, and only a poller that reads and compares it notices. - **Shortening the interval multiplies the load.** A 60-second cycle is five times the work: about 1,040 values a second for the same estate. ## What notifications give and what they miss A `linkDown` notification arrives within seconds of the event and costs nothing while nothing happens. Its gaps: - **A silent failure sends nothing.** A switch that loses power or hangs originates no notification; only a poll timeout reveals it. - **The path to the manager may be the thing that failed.** If a switch's only uplink goes down, its notification has no route to UDP 162. - **Loss.** A trap is a single unacknowledged UDP datagram and can be dropped; the informs that fix this belong to notification design, not to this choice. - **No rates.** Notifications report events, not the steady counter deltas that capacity and error trends need. - **Bursts.** One distribution-layer failure can make many neighbours report at once. ## Side by side | | Five-minute poll | Link notification | |---|---|---| | Detection delay | up to 300 s | seconds | | Steady load | ~208 values/s for this estate | near zero | | Catches a dead or hung switch | yes, by timeout | no | | Catches a flap between polls | only via `ifLastChange` | yes, if delivered | | Gives counter deltas | yes | no | | Survives the uplink it reports on | yes, the poll times out | no | ## Why estates run both - the original design RFC 1157 §3.2.3 states SNMP's intended strategy: monitoring of network state "is accomplished primarily by polling", and "a limited number of unsolicited messages (traps) guide the timing and focus of the polling". The two are complementary rather than competing: 1. Poll everything on a cycle you can afford, for counters and proof of life. 2. Accept notifications for state changes, and treat each as a trigger to poll that switch immediately to confirm and gather context. 3. Read `ifLastChange` and `sysUpTime` on each poll, so a flap or an agent restart that produced no delivered notification still shows up. 4. Keep the notification destination reachable over a path independent of the links being reported on where you can. The interval is the main dial: shorter polls buy freshness at a linear cost in load on the manager and on every agent's CPU, while notifications buy freshness for events only.
- How can an SNMP poller detect that a port flapped between two five-minute polls?Read IF-MIB's `ifLastChange` with `ifOperStatus` on every poll. It holds the `sysUpTime` value when the interface entered its current state, so if it moved since the previous poll while the status still reads `up`, the port left and re-entered that state in between. Compare against `sysUpTime` to rule out an agent restart.
- Why might the manager never receive a linkDown notification for a switch's uplink?The notification travels as a UDP datagram to the manager's port 162 over the network. If the failed uplink was the switch's only path, nothing can carry it; a trap is also unacknowledged, so it may simply be lost. Polling catches both cases, because the next request times out.
saying these in an interview costs you the question
- With traps configured, you can stop polling the switches altogether.
- A five-minute poll of ifOperStatus will catch every link flap.
- A switch that crashes will always send a notification before it goes.
- Polling every 10 seconds is free, so freshness costs nothing.
- Notifications can replace polling for utilisation graphs.