An SNMP SetRequest applies its variable bindings atomically, so why is that not enough to roll one VPN change out to forty routers, and what does NETCONF offer instead?
answer
- atomic, but over what scope?
- one PDU, one agent
- row creation spans several PDUs
- lock, stage, validate, commit
- a commit that reverts unless confirmed
basics
~20 sA SetRequest is atomic only within one PDU on one agent, while a fleet change spans many PDUs and devices with no lock, staging or rollback. NETCONF adds per-device locks, validated candidates and self-reverting confirmed commits that an orchestrator coordinates.
solid answer
~50 sRFC 3416 says the assignments in one `SetRequest-PDU` happen "as if simultaneously": if one fails, the agent undoes the rest and returns `commitFailed`. That atomicity covers one message to one agent. A VPN change needs many objects, often new table rows built through `RowStatus` over several PDUs, so a failure midway leaves a partial change for the manager to undo; SNMP has no protocol lock against other writers and no confirmed revert. RFC 3512 itself says relying on protocol-level transactional integrity is insufficient. NETCONF gives each device a `<lock>` that SNMP and CLI writes must also respect, a `candidate` datastore to stage and `<validate>`, an all-or-nothing `<commit>`, and a confirmed commit that reverts on its own. RFC 6241 Appendix E describes running those steps in parallel across devices, while calling complete multi-device transactions prohibitively expensive.
code
xml · 7 lines<rpc message-id="104"
xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
<commit>
<confirmed/>
<confirm-timeout>300</confirm-timeout>
</commit>
</rpc>go deeper
Remember that an SNMP SetRequest is atomic only for the bindings inside that one message, and that NETCONF lets you stage a change and commit it as a whole.
Explain the SetRequest phases and commitFailed, why row creation needs several PDUs, and the NETCONF sequence of lock, candidate edit, validate, commit and unlock.
Walk through a fleet rollout that fails on one router: what SNMP leaves behind, how a confirmed commit makes each router revert by itself, and why persist matters when the orchestrator restarts.
Be explicit that neither protocol gives a distributed transaction; argue for an orchestrator that stages and validates everywhere first and uses self-reverting commits to shrink the failure window.
## What SNMP actually promises An SNMP manager changes a device by sending a **`SetRequest-PDU`**: a list of **variable bindings**, each an object identifier and a new value. RFC 3416 (part of STD 62) defines two phases. First the agent validates every binding and rejects the request with an error such as `wrongValue`, `notWritable` or `inconsistentValue` before changing anything. Then it applies them, and "each of these variable assignments occurs as if simultaneously with respect to all other assignments specified in the same request". If an assignment still fails, the agent **undoes the others** and answers `commitFailed`; only if it cannot undo them does it answer `undoFailed`. So a single SetRequest is a small transaction. The trap is the **scope**: one message, one agent. ## Why a fleet change does not fit in that scope A VPN change on forty routers is many objects on many devices. - **It spans several PDUs on one device.** Message sizes are bounded, and new conceptual rows are usually created through a `RowStatus` column (RFC 2579), for example `createAndWait`, then the other columns, then `active`. Each step is a separate SetRequest, and nothing in the protocol ties them together. - **There is no lock.** SNMP has no lock operation; another manager or an operator at the CLI can change the same objects between your PDUs. A MIB module can build an advisory lock from a `TestAndIncr` object (RFC 2579), whose Set fails with `inconsistentValue` unless the manager supplies the current value, but only modules that define one have it. - **No staging or whole-configuration check.** Each PDU changes the live device; nothing validates the complete result before it takes effect. - **No automatic revert.** If the third PDU fails, the first two stay applied, and the manager must work out and send the inverse. - **Nothing spans devices.** Router 17 failing leaves routers 1-16 changed. The IETF says so itself. RFC 3535 notes that a logical operation "can turn into a sequence of SNMP interactions" where the implementation must keep state and roll back on failure, and RFC 3512 (Informational, on configuring with SNMP) concludes that "reliance on transactional integrity only at the SNMP protocol level is insufficient". ## What NETCONF provides per device RFC 6241 Appendix E.1 lists the steps for a safe change on one device, and each maps to a protocol primitive: 1. **Lock** the target datastore with `<lock>`. While held, other NETCONF sessions cannot edit it, and "SNMP and CLI requests to modify the resource MUST fail". 2. **Checkpoint** the running configuration. 3. **Load and validate** the change into the `candidate` datastore and run `<validate>`; YANG constraints are checked here. 4. **Change running** with `<commit>`, which applies the candidate as a whole. 5. **Test** the result, ideally under a **confirmed commit**: unless a confirming `<commit>` arrives within `confirm-timeout` (600 seconds by default), the device restores its previous configuration on its own. 6. **Make it permanent**, then **unlock**. The details of datastores, locks and confirmed commit belong to their own subjects; here the point is that they exist in the protocol rather than in each MIB module. ## Across forty routers Appendix E.2 describes the multi-device case: acquire locks on all targets, upload and validate everywhere, checkpoint, change running, test, and make permanent, restoring previous configurations wherever a step fails. It is explicit about the limit: > Providing complete transactional semantics across multiple devices is prohibitively expensive, but the size and number of windows for failure scenarios can be reduced. | Property | SNMP SetRequests | NETCONF | |---|---|---| | Atomic unit | one PDU on one agent | one `<commit>` of the candidate per device | | Lock against other writers | none in the protocol | `<lock>`, honoured by SNMP and CLI too | | Validate before applying | per binding only | `<validate>` of the whole candidate | | Automatic revert | within one PDU | confirmed commit, rollback-on-error | | Multi-device transaction | none | none built in; primitives for an orchestrator | The orchestrator still coordinates: stage and validate everywhere first, commit everywhere with a confirmed commit, test, then confirm. If it dies before confirming, each router reverts by itself, which shrinks the window for a half-changed network. ## The interview answer in one line SNMP's atomicity is real but scoped to one PDU on one agent; NETCONF supplies the per-device lock, staging, validation and self-reverting commit that make a network-wide change recoverable, without pretending to be a distributed transaction.
- Why would you add <persist> to a confirmed commit in a fleet rollout?Without `<persist>`, the confirming commit must come on the same session, and as soon as the server sees that session end it restores the old configuration. With `<persist>`, a later session can confirm or cancel by sending a matching `<persist-id>`, so a restarted orchestrator can finish the rollout instead of watching every router revert.
- What does SNMP return if an agent cannot undo a partly applied SetRequest?RFC 3416 says that if any assignment fails, the agent undoes the others and returns `commitFailed` with the failed binding's index. If and only if it cannot undo them all, it returns `undoFailed` with error-index zero, meaning the device's state is uncertain. Implementations are strongly encouraged to avoid both.
saying these in an interview costs you the question
- A SetRequest is atomic, so SNMP already gives you configuration transactions.
- If a later SetRequest fails, the agent undoes the earlier SetRequests too.
- NETCONF performs a true distributed transaction across all devices.
- A NETCONF lock only stops other NETCONF sessions, not SNMP or CLI writes.
- A confirmed commit needs the client to reconnect before the device can revert.