skip to content

An SNMP poller times out walking a 40,000-row table on a core router with chained GetNextRequests; how does GetBulkRequest fix it, and how do you size max-repetitions?

level: seniorimportance: must knowfreq 32%

answer

  1. round trips, not bytes
  2. N plus M times R
  3. trimmed from the end, not tooBig
  4. fit one unfragmented datagram
  5. small values under stress

basics

~20 s

Chained GetNext pays a round trip and a full message-processing pass per row. GetBulkRequest returns up to max-repetitions rows per exchange; size that value so each reply fits the manager's maximum message size and one unfragmented datagram, and keep it small under stress.

solid answer

~50 s

With GetNext, 40,000 rows cost 40,000 exchanges plus one to see the end; at 15 ms each that is about ten minutes, longer than a five-minute polling interval, and every timeout adds a retransmission. `GetBulkRequest` (RFC 3416) carries `non-repeaters` (N) and `max-repetitions` (M), and the reply holds up to N + M x R bindings, where R is the number of repeating bindings. Reading three columns with M = 10 returns ten rows per exchange: about 4,000 exchanges, roughly a minute. Size M so the reply fits the manager's maximum message size and, ideally, one unfragmented datagram: about 1,472 octets on a 1,500-octet MTU over IPv4. An oversized reply is trimmed from the end, not refused, and a fragmented one is lost whole if one fragment is lost; RFC 3416 says to use small values under network stress.

code

pseudocode · 14 lines
pseudocode
walk_column(agent, prefix, M):
  cursor = prefix
  loop:
    resp = GetBulk(agent, non_repeaters = 0, max_repetitions = M, names = [cursor])
    if resp.error_status != noError:
      fail(resp.error_status, resp.error_index)
    if resp.varbinds is empty:
      fail("no repetition fits the size limit: lower M")
    for vb in resp.varbinds:
      if vb.value is endOfMibView: return
      if not has_prefix(vb.name, prefix): return
      if vb.name <= cursor: fail("name did not increase")
      emit(vb.name, vb.value)
      cursor = vb.name

go deeper

for a junior

Remember that GetBulk asks for many successors in one exchange using non-repeaters and max-repetitions, and that this saves round trips over GetNext.

for a middle

Compute N + M x R for a request and explain why a GetBulk reply is trimmed from the end rather than refused with tooBig.

for a senior

Turn a timing-out walk into numbers: exchanges times round trip against the poll interval, then pick max-repetitions from the datagram budget and resume from the last name received.

for a principal

Set polling policy for a fleet: per-device request budgets, max-repetitions under congestion, retry limits, and when a table is too large for polling at the interval wanted.

## Why chained GetNext melts a poll A **GetNext walk** reads a table one row per exchange: the manager sends the last name it received and gets the next instance back. Each exchange costs a network round trip, and on the agent a complete pass through message processing: decoding the message, checking the community or the SNMPv3 security parameters (for authenticated SNMPv3, an HMAC computed on every message), checking the MIB view, finding the successor and encoding the reply. On many agents the successor lookup in a large table is itself expensive, but that depends on the implementation. Take the reserved scenario: a table of **40,000 rows** on a core router, three columns needed, one GetNext per row carrying one binding per column, and **15 ms** per exchange including agent processing. - Exchanges: 40,000 rows + 1 to discover the end = **40,001**. - Time: 40,001 x 0.015 s = about **600 s**, ten minutes. - If the poller collects every five minutes (300 s), the walk can never finish within the interval, and each timeout adds a retransmission to an agent that is already slow. RFC 3416 leaves retransmission to the manager but asks it to act responsibly (BCP 41); a poller that retries aggressively into a struggling agent multiplies the load it is trying to measure. ## What GetBulkRequest changes `GetBulkRequest-PDU` (tag `[5]`, sent in SNMPv2c and SNMPv3 messages) uses the same variable-binding list as GetNext, plus two fields in place of `error-status` and `error-index`: | Field | Meaning | |---|---| | `non-repeaters` (N) | the first N bindings get one successor each, like a GetNext | | `max-repetitions` (M) | each of the remaining R bindings gets up to M successors | The reply holds up to **N + (M x R)** bindings, ordered by repetition: the N non-repeaters, then the first successor of each of the R columns, then the second of each, and so on. RFC 3416's own example uses `non-repeaters = 1` for `sysUpTime`, which stamps every reply with the agent's uptime, useful when the walk collects counters for rate calculations. With N = 1 (`sysUpTime`), R = 3 columns and M = 10, each reply carries up to 1 + 30 = **31** bindings and **10 rows**: - Exchanges: 40,000 / 10 = **4,000** (one more may be needed to see the end). - Time: 4,000 x 0.015 s = about **60 s**, comfortably inside a five-minute interval. ## Sizing max-repetitions Bigger is not better. RFC 3416 bounds the reply by the smaller of the agent's local limit and the largest message the manager can accept, which an SNMPv3 manager advertises as `msgMaxSize`. A worked budget, with stated assumptions: 1. **Datagram budget.** 1,500-octet Ethernet MTU - 20-octet IPv4 header - 8-octet UDP header = **1,472 octets**, which is also the message size RFC 3417 recommends every SNMP implementation be able to accept over UDP (484 octets is the minimum). 2. **Assume about 100 octets** of message and PDU overhead and **about 40 octets per binding** (a long instance name plus a value); real sizes vary by object. 3. (1,472 - 100) / 40 = about 34 bindings: one for `sysUpTime` and 33 for **11 rows** of 3 columns. 4. Choose **M = 10** for headroom: 1 + 30 bindings is about 1,340 octets. Push M to 100 and a reply can approach 301 x 40 + 100 = 12,140 octets. If the manager accepts that, the datagram leaves as **nine** IPv4 fragments (each carries up to 1,480 octets of payload, and 12,148 / 1,480 = 8.2), and losing any one loses the reply; RFC 3416 calls fragmentation harmful and says that **under network stress only small values of max-repetitions should be used**. If the manager does not accept it, the agent trims the reply to fit, so the extra repetitions buy nothing on the wire. ## How an agent shortens a GetBulk reply RFC 3416 gives GetBulk **no `tooBig` path**. The reply may carry fewer bindings than N + M x R for three reasons: - it would exceed the size limits, so bindings are **removed from the end** until it fits; - every binding in some repetition is `endOfMibView`, so the rest can be cut; - the request needs far more processing than a normal one, so the agent stops early, provided at least one repetition is complete. Only if even an empty reply cannot be sent does the agent drop it silently and count it in `snmpSilentDrops`. The manager therefore resumes each column from the **last name it actually received**, never from where it hoped to be. ## Overshoot and the end of the table The last GetBulk reply usually runs past the end of the table into the next column or object; the manager discards bindings outside each column's prefix. With M = 10, that waste is at most one reply's worth of bindings, a small price for removing 36,000 round trips.

  • Why does a GetBulk walk resume from the last name received rather than counting M rows ahead?
    Because the agent may return fewer bindings than N + M x R: it trims from the end to fit the size limit, cuts after a repetition that is all endOfMibView, or stops early after at least one repetition when processing is expensive. Resuming from the last name actually received is correct for every one of those cases.
  • What happens if a GetBulk asks for max-repetitions far beyond the rows left in the table?
    The agent keeps returning successors past the table into the next column or object, or `endOfMibView` once nothing follows, and may truncate after a repetition that is entirely `endOfMibView`. Nothing fails; the manager discards bindings outside the column's prefix, and the cost is the wasted bytes and agent work.
  • Besides GetBulk, how can an SNMP poller shorten a long walk?
    Ask only for the columns it needs, put them in the same request so each reply covers a row of each, and spread work over several outstanding requests with distinct `request-id`s, for example one per column, while keeping retries modest. Collection intervals should comfortably exceed the measured walk time.

saying these in an interview costs you the question

  • A GetBulk reply that does not fit comes back with tooBig.
  • The largest max-repetitions value is always the fastest walk.
  • A GetBulk reply holds max-repetitions bindings in total.
  • GetNext walks are slow because each reply carries too many bytes.
  • The non-repeaters are repeated along with the other bindings.