skip to content

Locking and Transactions

Locks stop two automation systems interleaving half-applied changes, and a confirmed commit rolls back unless confirmed in time. Together they make NETCONF transactional where SNMP Set is not.

on this pageshow

questions

6

In NETCONF, what does a <lock> on a configuration datastore prevent, how long does it last, and what does a competing session get back?

level: middleimportance: must knowfreq 22%

answer

  1. one whole datastore per target
  2. changes blocked, from every interface
  3. ends at unlock or session end
  4. error-info names the holder

basics

~20 s

A NETCONF <lock> gives one session exclusive write access to a whole datastore until it unlocks or its session ends; other sessions, SNMP and CLI cannot change it, and a competing <lock> fails with lock-denied naming the holder's session-id.

solid answer

~50 s

`<lock>` takes a `<target>` naming one datastore - `running`, `candidate` or `startup` - and locks all of it. While it is held, the server must refuse changes from anyone else: another session's `<edit-config>` or `<copy-config>` into that datastore, and SNMP or CLI writes too (RFC 6241 §7.5). It guards changes, not reads. There is no lock timer: the lock lasts until the owner sends `<unlock>` or its session ends for any reason - `<close-session>`, a dropped transport, a server-side inactivity timeout, or another session's `<kill-session>`. A competing `<lock>` gets an `<rpc-error>` with `error-tag` `lock-denied` and the holder's `<session-id>` in `<error-info>`, or 0 when a non-NETCONF entity holds it. Only the owner can `<unlock>`. The server also refuses a lock on a candidate holding uncommitted edits, and on `running` while another session's confirmed commit is pending.

go deeper

for a junior

Remember that a NETCONF lock is taken per datastore, covers all of it, and is released by unlock or by the session ending.

for a middle

Explain what the lock blocks, including SNMP and CLI writes, why reads still work, and how lock-denied identifies the holder by session-id.

for a senior

Show you know a lock has no timer, so a dead client's lock lingers until the server notices; describe how a pipeline backs off or clears it safely.

for a principal

Weigh lock scope and hold time against operator access: one global lock serialises every change to a router, which is safe but slow when many systems share it.

## Why NETCONF has a lock Picture one router managed by two automation systems: a **CI pipeline** that pushes reviewed changes, and an **operator's script** that adjusts interface settings during an incident. Each sends a sequence of remote procedure calls - several `<edit-config>` requests, a check, perhaps a `<commit>`. Without coordination the two sequences interleave: the pipeline checks a configuration that the script changes a moment later, and the router ends up in a state that neither of them built or tested. The `<lock>` operation of **RFC 6241 §7.5** (NETCONF 1.1, which obsoletes RFC 4741) exists to stop that interleaving. A session takes the lock, makes its change and releases it. The RFC calls such locks **short-lived**: they protect one change, not a working day. ## What the lock covers `<lock>` has one mandatory parameter, `<target>`, which names a **configuration datastore**: `<running/>`, `<candidate/>` or `<startup/>`, each where the server supports it. The base lock is **global** to that datastore - it covers all of it, never a subtree. Locking only part of `running` is a separate extension, `<partial-lock>` from RFC 5717. On servers that implement the NMDA operations of RFC 8526, `<lock>` and `<unlock>` gain a `datastore` leaf, and only writable datastores can be locked. While the lock is held, the server must prevent any change to the datastore other than those the owning session requests: | Request from anyone other than the holder | Result while the lock is held | |---|---| | `<edit-config>` targeting the locked datastore | refused | | `<copy-config>` with the locked datastore as its target | refused | | An SNMP or CLI write to the locked configuration | MUST fail | | `<commit>` while `running` or `candidate` is locked by another session | fails with `in-use` | | A read such as `<get-config>` | allowed - the lock guards changes | | Another `<lock>` on the same datastore | fails with `lock-denied` | The SNMP and CLI row matters: a NETCONF lock is not just a courtesy between NETCONF clients. The RFC requires the device to enforce it against its other management interfaces too. ## How long a lock lasts The protocol has **no lock timer**. A lock begins when it is granted and ends in one of three ways: 1. The owner sends `<unlock>` for the same target. Only the session that obtained the lock may do this; an `<unlock>` from another session, or for a lock that is not active, fails. 2. The owning session ends. `<close-session>` releases its locks, and so does any other termination - the server may also close a session when the transport fails, after an inactivity timeout or on abusive behaviour, criteria RFC 6241 leaves to the implementation. 3. Another session ends the owner with `<kill-session>`, which releases every lock the killed session held. Because the end of a session is the only automatic release, a lock can outlive the work it protected for as long as the server still believes the dead client's session is alive. One datastore behaves specially: when a lock on `candidate` is released - explicitly or because the session failed - the **outstanding changes in the candidate are discarded** (RFC 6241 §8.3.5.2). A crashed client therefore does not leave half-made edits for the next committer. ## When the server refuses a lock A `<lock>` MUST NOT be granted when: - a lock on that datastore is already held by any NETCONF session or by another entity, such as a CLI user; - the target is `candidate` and it has been modified with changes that have been neither committed nor discarded; - the target is `running` and another NETCONF session has a **confirmed commit** still awaiting its confirmation; - on a server with RFC 5717, any partial lock exists on `running` - even one held by the requesting session itself. ## Reading lock-denied When the pipeline asks for the lock while the script holds it, the reply carries the holder's identity: ```xml <rpc-reply message-id="101" xmlns="urn:ietf:params:xml:ns:netconf:base:1.0"> <rpc-error> <error-type>protocol</error-type> <error-tag>lock-denied</error-tag> <error-severity>error</error-severity> <error-info> <session-id>454</session-id> </error-info> </rpc-error> </rpc-reply> ``` - `error-tag` **`lock-denied`** is the one RFC 6241 Appendix A defines for this case, with `error-type` `protocol`. - `<session-id>` is the NETCONF session that holds the lock - the same number that session received in its server `<hello>`. - A `<session-id>` of **0** means a **non-NETCONF entity** holds the lock, such as a CLI session. RFC 6241 defines no NETCONF way to break that lock and leaves it to the device's own interface. ## Operating notes - Take the lock **before the first edit**, not just before the commit; the point is that nobody changes the datastore between your first read and your last write. - On `lock-denied`, back off and retry, or investigate the holder; never fall back to editing unlocked. - Hold the lock for one change. A pipeline that keeps a lock across a slow approval step locks every operator out of the router. - A RESTCONF edit to a datastore that a NETCONF session has locked fails too, with `409 Conflict` and `in-use` (RFC 8040 §1.4).

  • Can the CI pipeline send <unlock> to clear a lock the operator's script took?
    No. RFC 6241 §7.6 lets only the session that obtained the lock release it; an `<unlock>` from any other session fails. The pipeline's options are to wait for the script's session to end, or to end that session itself with `<kill-session>`, using the `<session-id>` from the `lock-denied` reply.
  • Why might lock-denied report session-id 0?
    Zero means the lock is held by something that is not a NETCONF session - typically a CLI user who locked the configuration through the device's own interface. `<kill-session>` cannot reach it, because it has no NETCONF session to kill; RFC 6241 says breaking locks held by other entities is outside its scope.

saying these in an interview costs you the question

  • A NETCONF lock only binds other NETCONF clients, so CLI edits still get through.
  • The base NETCONF lock can cover a single subtree, such as one interface.
  • A NETCONF lock expires on its own after a fixed protocol timeout.
  • Any administrator's session can send <unlock> to clear another session's lock.
  • A lock on running stops other sessions from reading it with get-config.
open as a page

A CI pipeline's NETCONF commit to a router might cut off the pipeline's own management path; how does a confirmed commit protect it, and when does it roll back?

level: seniorimportance: must knowfreq 20%

basics

~20 s

A confirmed commit applies the candidate but reverts running to its prior state unless a confirming <commit> arrives within confirm-timeout (600 seconds by default); it also reverts at once if the issuing session drops, unless <persist> was set.

open as a page

Before a NETCONF client edits the candidate datastore and commits it, why should it lock both candidate and running, and what does each lock guard?

level: middleimportance: should knowfreq 12%

basics

~20 s

The candidate is a scratch pad every session may share, and <commit> publishes all of it; locking candidate keeps other sessions' edits out of your commit, and locking running stops anyone changing or committing to running until you finish.

open as a page

An operator's script died holding a NETCONF lock on running and the CI pipeline now gets lock-denied; how should the pipeline clear the stale lock safely?

level: seniorimportance: should knowfreq 10%

basics

~20 s

The pipeline takes the holder's session-id from lock-denied, confirms through NETCONF monitoring that the session is really stale, and sends <kill-session>; that releases the lock but does not roll back edits the dead session already made to running.

open as a page

Your NETCONF automation must push one change to forty routers while operators still run their own scripts; how would you design the locking and confirmed commits around it?

level: principalimportance: should knowfreq 7%

basics

~20 s

Lock running and candidate on each router, load and validate, then confirmed-commit everywhere, test the network, and confirm only when all pass; otherwise cancel. NETCONF gives per-device atomicity and an undo window, not a cross-device transaction.

open as a page

When would a NETCONF client use RFC 5717's <partial-lock> instead of <lock>, and what does a partial lock fail to protect?

level: seniorimportance: nice to knowfreq 5%

basics

~20 s

A <partial-lock> locks selected nodes of running so two managers can edit different sections at once; it protects only the nodes its XPath matched at lock time, plus their subtrees, never later-created siblings, reads, or the rest of the configuration.

open as a page