How would you divide configuration and monitoring between NETCONF and SNMP for a router fleet in which a third of the devices offer only SNMP?
answer
- one writer per device
- SNMP still wins at polling
- no write view for SNMP
- two event models, one timeline
- a trigger for retiring each path
basics
~20 sConfigure through NETCONF wherever devices support it, keep a single writer per device, and give SNMP read-only access; poll counters with SNMPv3 across the whole fleet, and accept traps or informs from legacy devices beside NETCONF notifications from the rest.
solid answer
~40 sI would split by job, not by vendor. Configuration goes through NETCONF on every device that supports it, using candidate, validate and confirmed commit, from one source of intended configuration; the SNMP-only third keeps whatever configuration path it has, with SNMP writes only where a standard writable MIB module really exists. Monitoring stays on SNMPv3 across the fleet, because RFC 3535 found it good at polling and `IF-MIB` is nearly universal; use the 64-bit counters on fast interfaces. Give SNMP managers a VACM access entry with no write view so SNMP cannot become a second writer outside NETCONF's locks. Events arrive as informs from legacy devices and RFC 5277 notifications from the rest, normalised onto one timeline. Name the trigger that retires each path, such as the last SNMP-only device leaving service.
go deeper
Recall the default split: NETCONF to configure, SNMP to poll counters, and why devices without NETCONF still need SNMP.
Explain how each half works in the hybrid: candidate, validate and confirmed commit for changes; SNMPv3 polling of IF-MIB 64-bit counters; informs versus RFC 5277 notifications for events.
Close the gaps between the halves: a NETCONF lock only protects while held, so remove SNMP write access with VACM, and normalise sysUpTime-relative and absolute event times.
Own the trade-off and its exit: two credential systems, two schemas and two event pipelines are a cost justified only by the legacy third, so define the trigger that retires each path.
## The situation An architecture review is choosing how to **configure** and **monitor** a fleet of routers. Two-thirds support NETCONF with YANG models; one-third offer SNMP and a command line only. There is no single right answer, but there is a defensible one, built from what each protocol was designed to do (RFC 3535, the 2002 IAB workshop) and from what goes wrong when both are allowed to write. ## Principles 1. **Split by job, not by device.** Monitoring and configuration have different needs: many small, frequent reads versus rare, validated, transactional writes. 2. **One writer per device.** Whatever configures a device should be the only thing that does. Two writers produce drift that neither notices. 3. **One source of intent.** Intended configuration lives in a central system; protocols are delivery mechanisms, not the record. 4. **Each compromise gets a retirement trigger,** not a date. ## Configuration | Device class | Configuration path | Why | |---|---|---| | NETCONF-capable | NETCONF: lock, edit `candidate`, `<validate>`, confirmed `<commit>` | validated, all-or-nothing per device, reverts on its own if not confirmed | | SNMP-only | the device's existing path; SNMP `SetRequest` only where a standard writable MIB module covers the feature | RFC 3535: standard modules seldom hold the writable objects you need | On NETCONF devices, a datastore `<lock>` makes SNMP and CLI writes to the locked resource fail, but only while it is held. Between changes, an SNMP manager with write access could still change things behind the automation's back. So: - Give SNMP managers a **View-based Access Control Model** (VACM, RFC 3415) access entry whose `vacmAccessWriteViewName` is empty; an empty name grants no write access. - Restrict NETCONF writes to the automation identity with the **NETCONF Access Control Model** (NACM, RFC 8341, which obsoletes RFC 6536). ## Monitoring SNMP keeps this job across the whole fleet, which is where it still wins: - **Ubiquity.** RFC 3535 found SNMP reasonable for monitoring and `IF-MIB` implemented on most devices; RFC 6632 notes that standard counters allow interoperable comparison across vendors. - **Uniform polling.** One poller, one schema of counters, old and new devices alike. - **Correct counters.** On fast interfaces poll the 64-bit `ifHC*` counters: a 32-bit octet counter wraps after 2^32 octets, about 34 seconds at 1 Gb/s and 3.4 seconds at 10 Gb/s. RFC 2863 requires 64-bit octet counters above 20 Mb/s and 64-bit packet counters too at 650 Mb/s and above. Check `ifCounterDiscontinuityTime` or `sysUpTime` so a restart is not read as traffic. - **Secure it.** Use SNMPv3 with authentication and privacy (the User-based Security Model, RFC 3414, or the Transport Security Model, RFC 5591, over TLS or DTLS, RFC 6353). SNMPv2c (RFC 1901, published as Experimental) authenticates only with plain-text community strings, and RFC 3410 records that it and SNMPv1 were declared Historic once SNMPv3 became a full Standard. NETCONF's `<get>` can read state on the newer devices, and that is useful for verifying a change against its intent. It does not need to replace the poller for routine counters. ## Events - From SNMP-only devices: prefer **informs** over traps, since an `InformRequest-PDU` is acknowledged and retransmitted. - From NETCONF devices: RFC 5277 **notifications** on a dedicated subscription session, resubscribing with a `<startTime>` after a reconnect where replay is supported. - **Normalise time.** SNMP notifications carry `sysUpTime.0`, relative to an agent restart; NETCONF carries an absolute `<eventTime>`. Convert both onto one timeline before correlating. Streamed telemetry is another option for the newer devices, but it is a separate decision with its own trade-offs. ## What the hybrid costs - **Two credential systems** (SNMPv3 users and keys, SSH identities), which RFC 3535 already flagged as a burden for SNMP. - **Two schemas** to map onto one inventory: MIB objects and YANG nodes. - **Two event pipelines** with different loss models. Those costs are worth paying only while the legacy third exists. The retirement trigger for SNMP writes is the last device that needs them leaving service; SNMP polling can stay as long as it is the cheapest uniform view of counters.
- Why not configure the NETCONF-capable devices through SNMP too, so there is one configuration path?Because uniformity at the lowest common denominator throws away what you need: whole-configuration retrieval, validation before applying, an all-or-nothing commit and a self-reverting confirmed commit. RFC 3535 also found standard MIB modules seldom expose the writable objects a router needs, so the single path would still fall back to the device's own interface for most features.
- A team proposes moving all monitoring to NETCONF <get> on the newer devices; what would you ask first?What it gains over the existing poller. `<get>` returns modelled state, which helps verify a change against intent, but routine counter polling across a mixed fleet is where SNMP is uniform and cheap. I would ask about load on the devices, how counters map between MIB objects and YANG nodes, and whether two monitoring paths for one metric are worth running.
saying these in an interview costs you the question
- Running both NETCONF and SNMP is always a design smell; pick one.
- Let SNMP and NETCONF both write configuration so either can fix a device.
- A NETCONF lock protects the device from SNMP writes at all times.
- SNMPv2c communities are fine because the SNMP side is read-only.
- Polling 32-bit octet counters every five minutes is fine on 10 Gb/s links.