skip to content

An operator's script died holding a NETCONF lock on running and the CI pipeline now gets lock-denied; how should the pipeline clear the stale lock safely?

level: seniorimportance: should knowfreq 10%

answer

  1. the error names a session
  2. check the holder before acting
  3. kill-session releases, never undoes
  4. zero means not NETCONF
  5. denied by default under NACM

basics

~20 s

The pipeline takes the holder's session-id from lock-denied, confirms through NETCONF monitoring that the session is really stale, and sends <kill-session>; that releases the lock but does not roll back edits the dead session already made to running.

solid answer

~50 s

The `lock-denied` reply carries the holder's `<session-id>`. A lock lasts until its session ends, and a server decides a session has ended on criteria RFC 6241 leaves to the implementation, so a dead script's lock can linger. Before acting, the pipeline checks the holder: the RFC 6022 monitoring data lists each session's username, source host and login time, and each datastore's lock with `locked-by-session` and `locked-time`. If it is stale, `<kill-session>` with that id makes the server abort the session's operations, release its locks and close its connection. It does **not** roll back changes the session already made, except that a confirmed commit in progress is restored. So inspect `running` afterwards. A `session-id` of 0 is a non-NETCONF holder that `<kill-session>` cannot reach, and under RFC 8341 access control `<kill-session>` is denied unless a rule allows it.

go deeper

for a junior

Remember that lock-denied names the holding session, and kill-session ends that session and releases its locks.

for a middle

Explain why a dead client's lock can linger, how monitoring data identifies the holder, and what session-id 0 means.

for a senior

Stress that kill-session does not undo edits already applied to running, so the pipeline reconciles afterwards; note the confirmed-commit exception and the RFC 8341 default denial.

for a principal

Set the policy: who may break whose locks, after how long, and when the pipeline should page a human instead of forcing its way through.

## Why a lock goes stale A NETCONF lock has no timer. RFC 6241 §7.5 says it lasts until it is released or the **session closes**, and that the server may close a session itself on criteria such as transport failure, an inactivity timeout or abusive behaviour - criteria that depend on the implementation and the transport. So when an operator's script crashes, or its host loses power, the router may keep the session, and its lock, for as long as it takes to notice. Meanwhile the **CI pipeline** wants to deploy. Its `<lock>` on `running` comes back with `error-tag` `lock-denied` and, in `<error-info>`, the holder's `<session-id>` - say 454. ## Step 1: identify the holder before acting The session-id alone does not say whether the holder is dead or merely slow. The NETCONF monitoring module (**RFC 6022**, `ietf-netconf-monitoring`) answers that from the device itself: | Where | What it shows | |---|---| | `/netconf-state/sessions` | each session's `session-id`, `username`, `source-host`, `login-time` | | `/netconf-state/datastores/datastore/locks` | a `global-lock` or `partial-lock` entry with `locked-by-session` and `locked-time` | A lock taken long ago by a session from a host that no longer runs the script is stale. A lock taken seconds ago is a colleague mid-change, and the right answer is to back off and retry. ## Step 2: kill-session, and what it does `<kill-session>` (RFC 6241 §7.9) takes one parameter, the `session-id` to terminate. The server: 1. aborts any operations that session has in progress; 2. releases every lock and resource it holds; 3. closes its connections. Edge cases worth knowing: - Passing **your own** session-id returns `invalid-value`; a session leaves with `<close-session>` instead. - Releasing a killed session's lock on `candidate` **discards its outstanding candidate changes** (RFC 6241 §8.3.5.2). - RFC 6470's `netconf-session-end` notification reports the termination with `termination-reason` `killed` and a `killed-by` session-id - the audit record of who broke whose lock. ## Step 3: check what the dead session left behind This is the part weak answers miss. Apart from one case, `<kill-session>` **does not roll back** configuration changes made by the session that held the lock. If the script was editing `running` directly on a `:writable-running` device and died after three of five `<edit-config>` requests, those three edits stay. The exception: if the server receives `<kill-session>` while a **confirmed commit** is in progress, it MUST restore the configuration to its state before that confirmed commit. So after the kill, the pipeline takes its own lock and **reads `running`** before building on it. A script that worked through the candidate and died before committing is the safer case: its candidate edits were discarded with its lock, and `running` was never touched. | What the dead session had done | After `<kill-session>` | |---|---| | Uncommitted edits in a locked candidate | discarded with the lock | | `<edit-config>` requests already applied to `running` | **left in place** | | A confirmed commit still awaiting confirmation | restored to the pre-commit state | ## When kill-session cannot help - **`session-id` 0** in `lock-denied` means a non-NETCONF entity, such as a CLI user, holds the lock. There is no NETCONF session to kill; RFC 6241 says breaking such locks is beyond its scope, which in practice means the device's own interface. - **Access control.** Under the NETCONF Access Control Model (**RFC 8341**), access to `<kill-session>` is **denied by default**, and the `exec-default` leaf does not apply to it. The pipeline's account needs an explicit rule. That default is deliberate: `<kill-session>` changes no datastore, but it lets one session disrupt another's edit. - **Races.** RFC 6241's security section says breaking a lock by killing the session suffers from a race condition, which it suggests easing by removing the offending user from the AAA server. And killing a live session mid-edit leaves `running` as half-applied as a crash would. ## A safe policy for automation - Retry with back-off first; most `lock-denied` replies are a colleague's short change. - Kill only when monitoring shows the holder is stale by a threshold the team agreed in advance. - Never kill on `session-id` 0; escalate to whoever owns the CLI session. - After a kill, read `running` and reconcile before changing anything. - Give `<kill-session>` rights to the pipeline's account only if the team wants the pipeline to win; otherwise page a human.

  • Why not let the pipeline kill any session that blocks it?
    Because a lock is usually a colleague's short, legitimate change. Killing a live session aborts its operations without undoing what it already wrote, so running can be left half-changed - worse than waiting. Kill only when monitoring shows the holder is stale, and read running before building on it.
  • What does the operator's dead session leave behind if it was working through the candidate?
    Nothing in running. Its edits sat in the candidate, and RFC 6241 §8.3.5.2 discards outstanding candidate changes when the candidate lock is released - including when kill-session releases it. Working through a locked candidate is what makes a crashed client cheap to clean up.

saying these in an interview costs you the question

  • kill-session rolls back every change the killed session made to running.
  • A stale NETCONF lock expires by itself after a protocol-defined timeout.
  • kill-session with session-id 0 clears a lock held from the CLI.
  • kill-session is harmless, so any authenticated user is allowed to send it.
  • The pipeline can clear the lock by sending <unlock> on the dead session's behalf.