skip to content

Your NETCONF automation must push one change to forty routers while operators still run their own scripts; how would you design the locking and confirmed commits around it?

level: principalimportance: should knowfreq 7%

answer

  1. per device: lock, commit, confirm
  2. no cross-device transaction
  3. the confirm phase is the commit point
  4. timer long enough, or extend it
  5. persist trades one safety for another

basics

~20 s

Lock running and candidate on each router, load and validate, then confirmed-commit everywhere, test the network, and confirm only when all pass; otherwise cancel. NETCONF gives per-device atomicity and an undo window, not a cross-device transaction.

solid answer

~50 s

NETCONF has no multi-device transaction; RFC 6241 Appendix E calls full transactional semantics across devices prohibitively expensive, but its primitives shrink the failure windows. On each router: lock `running` and `candidate`, edit and validate the candidate, then send a confirmed commit with a timeout longer than the whole roll-out plus the network-wide tests, or extend it with follow-up confirmed commits. Confirm everywhere only when every router passes; on any failure, `<cancel-commit>` on every router or let the timers expire. The confirm phase is the real commit point: if some routers confirm and the rest revert, the fleet is split. Decide whether to use `<persist>` - it survives a controller restart but loses the revert-on-drop that catches a lockout. Hold locks for minutes, not hours, so operators are not locked out, and design for lock contention with back-off and a policy on who may kill whose session.

go deeper

for a junior

Remember that NETCONF changes one device at a time, and a confirmed commit reverts each device on its own.

for a middle

Lay out the per-device sequence of lock, edit, validate, confirmed commit, confirm and unlock, and say what each step protects.

for a senior

Explain why the confirm phase is the dangerous window, how follow-up confirmed commits extend timers, and how locks and back-off handle operators' scripts.

for a principal

Classify the change first, then justify lock scope, hold time, timeout and persist as stated trade-offs, naming the failure you accept.

## What NETCONF gives you, and what it does not On one router, NETCONF can make a change close to transactional: locks keep other writers out, `<commit>` from a locked candidate is all-or-nothing (RFC 6241 §8.3.4.1), and a **confirmed commit** reverts unless confirmed (§8.4). Across forty routers there is no such primitive. RFC 6241 Appendix E.2 (non-normative) says so directly: full transactional semantics across devices are **prohibitively expensive**, but the protocol has enough primitives to **reduce the size and number of failure windows**. Appendix E.2 also splits fleet changes into two classes, and the design starts by deciding which one this is: - **Independent:** a failure on one router can be retried or reported - adding an NTP server. - **All-or-nothing:** the network must reach the new state or return to the old one - a change several routers must agree on. ## A per-device skeleton For each router the controller runs the same steps (Appendix E.1): 1. `<lock>` `running`, then `candidate`; on `lock-denied`, back off. 2. `<edit-config>` into `candidate`, then `<validate>`. 3. `<commit>` with `<confirmed/>` and a chosen `<confirm-timeout>`. 4. Hold here until the fleet-wide decision. 5. Confirming `<commit>` - or `<cancel-commit>`. 6. With `:startup`, `<copy-config>` from `running` to `startup`. 7. `<unlock>` `candidate` and `running`. For an all-or-nothing change, Appendix E.2 has locks taken on every device and kept until every device is updated and the change made permanent. ## The decisions that have no single right answer | Decision | Option A | Option B | What it costs | |---|---|---|---| | Lock scope | global `<lock>` on each router | RFC 5717 `<partial-lock>` on the affected subtree | global locks serialise operators; partial locks work on `running` only and miss nodes created later | | Lock hold time | lock across the whole fleet roll-out | lock per router, release after its confirm | long holds lock operators out; short holds let another change land between routers | | `confirm-timeout` | long enough for the full roll-out and tests | short, extended with follow-up confirmed commits | a long window leaves a bad change live longer; extensions add calls that can fail | | `<persist>` | set it | leave it out | with it, a controller restart does not revert forty routers; without it, a session dropped by a lockout reverts at once | ## The real commit point Every router is provisional until confirmed. The dangerous moment is not the commits but the **confirms**: if the controller confirms twenty routers and then crashes, the other twenty revert when their timers expire, and the fleet is split between the old and new configurations. Ways to narrow that window: - send the confirming commits in a fast parallel burst, after all tests pass; - make the timeout comfortably longer than that burst; - with `<persist>`, a restarted controller can finish the confirms from new sessions by quoting each `<persist-id>`; - watch RFC 6470's `netconf-confirmed-commit` notifications, whose `confirm-event` values `start`, `cancel`, `timeout`, `extend` and `complete` show each router's state. A follow-up confirmed commit resets a router's timer to the new value, which is how a slow roll-out keeps early routers from reverting before the last one is ready. ## Living with operators' scripts - **Contention:** while a confirmed commit is pending, another session cannot lock `running` (RFC 6241 §7.5), and a partial lock is denied with `outstanding-confirmed-commit` (RFC 5717). Operators' scripts must treat those refusals as back-off signals. - **Ordering:** if two automation systems lock routers in different orders, each can hold half the fleet and wait on the other. A fixed lock order, or releasing everything and retrying after a randomised wait, avoids that. RFC 5717 gives the back-off rule for partial locks within one device; extending it to a fleet is a design choice. - **Breaking locks:** `<kill-session>` is denied by default under RFC 8341, does not roll back finished edits, and cannot reach a CLI holder (`session-id` 0). Decide in advance which system may break whose locks. - **Reverts undo others' work:** a revert restores the configuration from before the confirmed commit, so anything another session changed meanwhile is at risk unless locks kept it out. ## What a principal answer weighs There is no correct setting, only a stated trade-off: safety against lock-out time, automatic revert against controller restarts, and a whole-fleet window against a split fleet. A strong answer names the change class first, then picks lock scope, hold time, timeout and `persist` to match it, and says which failure it accepts.

  • Is a confirmed commit on every router a two-phase commit across the fleet?
    Only loosely. It gives each router a provisional state and an undo, like a prepare phase, but there is no coordinator protocol: each router reverts on its own timer, and nothing stops some routers confirming while others revert. The confirm burst is the commit point, and the design must narrow that window rather than assume atomicity.
  • Why might you release each router's lock as soon as it confirms instead of holding all forty?
    Holding every lock for the whole roll-out keeps operators off forty routers for its full length. Releasing per router frees them sooner, but lets another change land between routers. For an independent change that is usually acceptable; for an all-or-nothing change, Appendix E.2 keeps the locks until every device is done.

saying these in an interview costs you the question

  • NETCONF's confirmed-commit capability makes a change atomic across many devices.
  • One <lock> request can lock running on all forty routers at once.
  • Longer confirm-timeouts are always safer, so set them to hours.
  • persist is always better, because it never loses a change.
  • Once every router has replied ok to its confirmed commit, the change is permanent.