skip to content

An SNMP utilisation graph for a 10 Gb/s uplink drops to zero or spikes absurdly after a line card swap; how does IF-MIB let a poller detect counter discontinuities?

level: seniorimportance: should knowfreq 20%

answer

  1. counters have no starting value
  2. two clocks to compare
  3. the agent's uptime
  4. a per-interface timestamp
  5. discard, never interpolate

basics

~20 s

IF-MIB counters can restart at agent re-initialisation and at times marked by ifCounterDiscontinuityTime. A poller must discard any delta where sysUpTime went backwards or ifCounterDiscontinuityTime changed between polls, instead of treating the drop as a wrap.

solid answer

~40 s

A counter's absolute value has no meaning, and RFC 2863 says its value can jump at two kinds of event: re-initialisation of the agent's management system, which `sysUpTime` going backwards reveals, and other events such as a removable card being replaced, recorded per interface in `ifCounterDiscontinuityTime` (a `TimeStamp`: the `sysUpTime` value at the most recent discontinuity, zero if none). The RFC requires a manager to discard any difference for which `ifCounterDiscontinuityTime` differs between the two polls, in addition to the `sysUpTime` check. A naive poller does the opposite: it treats the smaller second reading as a wrap and adds 2^64, drawing an absurd spike, or clamps a negative delta to zero, drawing a fake outage. Also re-check the `ifIndex` mapping after a re-initialisation, because RFC 2863 only keeps `ifIndex` constant between re-initialisations.

code

pseudocode · 13 lines
pseudocode
new = get(sysUpTime.0, ifCounterDiscontinuityTime.i, ifHCInOctets.i)
if last is empty:
    last = new; return no_sample
if new.sysUpTime < last.sysUpTime:
    remap_ifindex(); last = new; return gap        # agent re-initialised
if new.ifCounterDiscontinuityTime != last.ifCounterDiscontinuityTime:
    last = new; return gap                          # this interface's counters restarted
if new.ifHCInOctets < last.ifHCInOctets:
    last = new; return gap                          # a Counter64 never wraps in practice
delta = new.ifHCInOctets - last.ifHCInOctets
seconds = (new.sysUpTime - last.sysUpTime) / 100
last = new
return delta * 8 / seconds                          # bit/s

go deeper

for a junior

Recall that SNMP counters can restart, for example after a reboot, and that a poller must not compute traffic across a restart.

for a middle

Explain the two signals, sysUpTime going backwards and ifCounterDiscontinuityTime changing, and why each covers a different kind of restart.

for a senior

Diagnose a spike or a false zero: tell a restart from a wrap, discard and record a gap, re-map ifIndex after re-initialisation, and use ifOperStatus to rule out a real outage.

for a principal

Decide how a monitoring platform treats gaps in capacity reports and billing-grade counts, and how much trust to place in agents that implement discontinuity signalling badly.

## Why a counter can jump SMIv2 (RFC 2578) says counters have no defined initial value and that discontinuities "normally occur at re-initialization of the management system, and at other times as specified in the description of an object-type". Every IF-MIB counter (RFC 2863), 32-bit or 64-bit, carries the same sentence: discontinuities can occur at re-initialisation "and at other times as indicated by the value of `ifCounterDiscontinuityTime`". The two kinds of event need two signals: | Event | What happens to counters | Signal a poller reads | |---|---|---| | Agent re-initialisation (a reboot, a restart of the management system) | all counters restart; `ifIndex` values may be reassigned | `sysUpTime.0` is smaller than at the previous poll | | An interface's counters restart without an agent restart (for example counters kept on removable hardware that was replaced) | that interface's counters restart, `ifIndex` is kept | that row's `ifCounterDiscontinuityTime` has changed | `ifCounterDiscontinuityTime` sits in `ifXTable`. Its value is the `sysUpTime` at the most recent discontinuity of any `Counter32` or `Counter64` in that interface's `ifTable` or `ifXTable` row, or zero if there has been none since the last re-initialisation. RFC 2863 introduced it so that an agent could keep the same `ifIndex` for a returning interface instead of having to assign a new one, and it asks agents not to update it unless absolutely necessary, because every update costs the manager a discarded sample. ## The rule a poller must follow RFC 2863 section 3.1.5 is explicit: a management application calculating differences between successive polls "must discard any calculated difference for which the value of `ifCounterDiscontinuityTime` is different for the two polls", and this is "in addition to the normal checking of `sysUpTime`". A correct poll cycle for one interface: 1. In one request, read `sysUpTime.0`, `ifCounterDiscontinuityTime.i` and the counters for interface `i`. 2. If `sysUpTime` is lower than last time, the agent re-initialised: discard the delta, re-learn which `ifIndex` is which interface, and start afresh. 3. If `ifCounterDiscontinuityTime.i` differs from last time, discard the delta for this interface. 4. Otherwise compute the delta, applying wrap correction only to a `Counter32` and only when at most one wrap was possible. 5. Store the new readings as the baseline for the next poll. The discarded interval should be stored as a gap. Interpolating across it, or writing zero, invents data. ## How the two symptoms arise - **The absurd spike**: the counter restarted, the new reading is smaller, and the poller assumes a wrap. With a `Counter64` it adds 2^64 - old + new, a number that implies more traffic than the link could carry in centuries. At 10 Gb/s a real `Counter64` wrap needs about 468 years, so on a 64-bit counter a backwards step is a restart, never a wrap. - **The drop to zero**: the poller computes a negative delta and clamps it to zero, or stores the small post-restart delta, or it follows an `ifIndex` that now belongs to a different, quiet interface after the reboot. - **A real zero**: the interface may genuinely be down. `ifOperStatus` and `ifLastChange` from the same poll separate a real outage from a counter artefact. ## Edge cases worth knowing - `ifIndex` must stay constant only "from one re-initialization of the entity's network management system to the next". After a reboot a poller keyed on `ifIndex` alone can silently graph the wrong port; matching on `ifName`, `ifDescr` or `ifAlias` re-establishes the mapping. - `sysUpTime` is `TimeTicks`, hundredths of a second modulo 2^32, so it wraps after about 497 days. A smaller `sysUpTime` on a device with very long uptime can be that wrap rather than a reboot, which is one reason to also watch `ifCounterDiscontinuityTime`. - Many agents and managers historically implemented this poorly, as the RFC itself admits, so a poller should treat the checks as mandatory rather than trust that discontinuities are rare.

  • Why does RFC 2863 ask agents not to update ifCounterDiscontinuityTime unless absolutely necessary?
    Every change forces the manager to throw away that interval's delta for the interface, which wastes polling and leaves a gap in the data. The object exists so an agent can keep a returning interface's `ifIndex` while still telling managers its counters restarted; it is not a reset button to press casually.
  • A device has run for 500 days and sysUpTime suddenly reads lower; did it reboot?
    Not necessarily. `sysUpTime` is `TimeTicks`, hundredths of a second modulo 2^32, which wraps after about 497 days. A poller should cross-check, for example whether the counters also restarted or whether `ifCounterDiscontinuityTime` and the `ifIndex` mapping changed, before deciding the agent re-initialised.

saying these in an interview costs you the question

  • Whenever the second reading is smaller, add 2^64 and carry on
  • Checking sysUpTime alone catches every counter discontinuity
  • ifIndex is guaranteed stable across reboots, so mappings never need re-learning
  • A counter discontinuity is the same thing as a counter wrap
  • Fill the gap after a restart by interpolating between neighbours
  • A graph at zero always means the interface carried no traffic