skip to content

An SNMP agent sends a core link-down as an InformRequest while its collector restarts for 100 seconds; with the SNMP-TARGET-MIB default timeout and retry count, is it delivered?

level: seniorimportance: should knowfreq 14%

answer

  1. TimeInterval counts hundredths
  2. 1500 and 3 are DEFVALs
  3. attempts = retries + 1
  4. the wait may be derived

basics

~20 s

Not with a fixed wait: the defaults are a 15-second timeout (1500 hundredths) and 3 retries, so four sends at 0, 15, 30 and 45 s, abandoned at 60 s, 40 s before the collector returns.

solid answer

~50 s

`snmpTargetAddrTimeout` defaults to 1500 and its type, TimeInterval, counts hundredths of a second, so 15 s; `snmpTargetAddrRetryCount` defaults to 3. One send plus three retries, with a fixed 15 s wait, puts transmissions at 0, 15, 30 and 45 s, and the agent declares failure at 60 s. The collector is back at 100 s, so the link-down is lost — and so is any inform raised in the first 55 s of the outage. RFC 3413 lets an implementation derive the real wait, for example by backing off; with doubling waits the sends fall at 0, 15, 45 and 105 s, and the fourth one lands. To survive a 100 s gap with a fixed 15 s timeout you need at least 7 retries, which means each unanswered inform is held for 120 s — memory the agent pays for every pending notification.

go deeper

for a junior

Recall that an inform is resent a limited number of times and then dropped, so a long collector outage still loses it.

for a middle

Explain the two SNMP-TARGET-MIB objects, their units and defaults, and compute the attempt times and give-up time from them.

for a senior

Size retries to the longest planned outage, account for the agent memory each pending inform holds, and check whether the device backs off before trusting the arithmetic.

for a principal

Decide how much outage the devices should absorb with retries versus a second receiver or reconciliation, given state per device and the burst on recovery.

## The objects and their units An agent that sends notifications as informs takes its timing from the **SNMP-TARGET-MIB** (RFC 3413), one row per destination in `snmpTargetAddrTable`: - **`snmpTargetAddrTimeout`** — type `TimeInterval`, which RFC 2579 defines as "a period of time, measured in units of 0.01 seconds". `DEFVAL { 1500 }`, so **15 seconds**. - **`snmpTargetAddrRetryCount`** — `Integer32 (0..255)`, `DEFVAL { 3 }`. - **`snmpNotifyType`**, in the companion SNMP-NOTIFICATION-MIB — `trap(1)` or `inform(2)`; only `inform(2)` uses the two values above. Two caveats from the RFC itself. The DEFVALs are what a row gets when it is created without those columns; a device's own configuration may set other values. And the timeout is "the expected maximum round trip time": RFC 3413 says the actual wait "may actually be derived from the value of this object", for example by a retransmission algorithm that depends on how many timeouts have already occurred — the method is implementation dependent. An application may also supply its own retry count. ## The timeline with a fixed wait Retry count 3 means **four** transmissions: the original and three retries. Suppose the core link fails at t = 0, the instant the collector stops, and the collector is listening again at t = 100 s. | Event | Time (s) | Collector | |---|---|---| | Original InformRequest | 0 | down | | Retry 1 | 15 | down | | Retry 2 | 30 | down | | Retry 3 | 45 | down | | Agent gives up | 60 | down | | Collector back | 100 | up | The agent abandons the notification **40 seconds** before anyone could have received it. Generalise: an event raised at time `e` has its last send at `e + 45`, so with a 100 s outage every inform raised before **t = 55 s** is lost. A trap raised at any point in the outage is lost too — the difference is that the agent knows about the informs. ## The same defaults with back-off If an implementation doubles the wait after each timeout (one derivation the RFC permits, not one it requires), the waits are 15, 30, 60 and 120 s: 1. send at 0, wait 15; 2. retry at 15, wait 30; 3. retry at 45, wait 60; 4. retry at **105** — after the collector returned — and it is acknowledged. So the honest answer to "is it delivered?" is **no with a fixed wait, yes with this back-off** — and the operator must find out which the device does rather than assume. ## Sizing the retries to the outage With a fixed wait `T` and retry count `R`, the last send is at `R × T` and the agent holds the inform for `(R + 1) × T`. To have a send land after an outage of length `D`, you need `R × T > D`: - `T` = 15 s, `D` = 100 s: `R × 15 > 100` gives **R ≥ 7** (6 retries end at 90 s, still inside the outage). The last send is at 105 s. - Each unanswered inform is then held for `8 × 15` = **120 s**. - At 20 new informs per second towards that dead collector — a plausible burst when a core link takes neighbours with it — the agent holds `20 × 120` = **2,400** pending informs for that one target, and twice that if a second receiver is down too. Longer timeouts or more retries cover longer outages, but they move the cost onto the device: state per pending inform, retransmissions into a dead port, and a burst of late arrivals when the collector comes back. ## What this means for the design - **Know your outage budget.** A collector restart, an upgrade or a failover has a duration; the inform retry window should exceed it, or a second receiver should cover it. - **Do not rely on the retry window alone.** Any finite window loses events from an outage that outlasts it — which is why RFC 3416 says an inform has no guarantee of delivery. - **Reconcile after the gap.** After a collector restart, polling `ifLastChange` and `sysUpTime` on the devices finds transitions that no notification reported. ## The interview point The strong answer reads the units correctly (hundredths, not milliseconds), counts attempts as retries plus one, notices that the RFC leaves the wait schedule to the implementation, and turns the arithmetic into a design rule: size the retry window to the longest planned outage and still reconcile afterwards.

  • Why not simply set the retry count to its maximum of 255?
    Because every pending inform is state on the agent and traffic on the network. With a fixed 15 s wait, 255 retries hold each notification for 256 × 15 = 3,840 s, over an hour, resending it every 15 s into a dead port. In an event storm that state grows with the event rate, and the returning collector gets the backlog at once. Size retries to planned outages; cover longer ones with a second receiver and reconciliation.
  • Would sending the same notification as a trap during the outage have been any worse?
    For delivery, no: both the trap and the defaults-timed inform are lost. The difference is knowledge. The inform's originator knows the acknowledgement failed and can record it; the trap's originator has no idea. That knowledge is only useful if something reads it, so it must be exported or polled, and a reconciliation poll after the gap is still needed.

saying these in an interview costs you the question

  • snmpTargetAddrTimeout 1500 means 1,500 milliseconds.
  • A retry count of 3 means three transmissions in total.
  • The RFC fixes a constant wait between inform retransmissions.
  • Default inform settings cover any normal collector restart.
  • Raising the retry count has no cost on the device.
  • Once retries are exhausted the agent switches to sending a trap.