Would you build a one-VLAN self-service portal and a nightly network-wide change window on RESTCONF, NETCONF or both, and why?
answer
- match the protocol to the unit of change
- web stack versus session client
- commit is per device
- narrow the failure window, don't close it
basics
~20 sRESTCONF fits the portal: small independent edits from web tooling, with entity-tag checks. NETCONF fits the window: locks, candidate, validate and confirmed commit per device. Neither makes a multi-device change atomic; the orchestrator owns that.
solid answer
~40 sI would argue it per workflow. The portal makes one small change on one device per click, from a web stack: RESTCONF's HTTPS, JSON and per-request edits fit, `If-Match` stops a stale overwrite, and YANG Patch covers a multi-node edit. The nightly window changes many things on many devices and must be safe to stop halfway: NETCONF lets the orchestrator lock each device, stage and validate in the candidate, confirm-commit everywhere, check, then confirm. Neither protocol gives atomicity across devices — RFC 6241 Appendix E calls complete multi-device transactional semantics "prohibitively expensive" — so a device that fails or loses reachability leaves a partial network the orchestrator must reconcile. If both run against the same devices, the window must lock, or portal edits will commit or confirm its half-finished work.
go deeper
Remember the short version: RESTCONF suits small edits from web applications, NETCONF suits large staged changes.
Explain what each workflow needs, per-request edits and entity-tags versus locks, candidate and commit, and how each protocol supplies it.
Walk through a multi-device window with locks, validation and confirmed commits, and the partial outcomes it can still leave behind.
Argue the trade-off rather than a winner: tooling and identity cost versus transactional control, the orchestrator's ownership of cross-device consistency, and the rule both writers must follow on shared devices.
## Two workflows with different units of change The portal and the change window look similar — both push YANG-modelled configuration — but they differ in everything that decides the protocol: | | Self-service portal | Nightly change window | |---|---|---| | Size of a change | one VLAN, a few nodes | many modules per device | | Devices per change | usually one | many | | Initiator | many users, any time | one orchestrator, scheduled | | Cost of a half-applied change | small, local | outages, inconsistent policy | | Client software | web application | automation engine | There is no single winner; there is a fit. ## The case for RESTCONF in the portal - **Tooling.** RESTCONF is HTTP over TLS with XML or JSON bodies (RFC 8040), so the portal's own HTTP stack, proxies and JSON handling work unchanged. - **Authentication.** RFC 8040 §2.5 prefers TLS client certificates and allows a registered HTTP authentication scheme, which fits a web service calling devices with its own credentials. NETCONF's mandatory SSH mapping means SSH accounts or keys, or mutual X.509 under RFC 7589. - **The change is already atomic enough.** One VLAN edit is one request. When it touches several nodes, a YANG Patch (RFC 8072) applies the ordered edits all-or-nothing. - **Concurrency.** Two users editing the same VLAN are caught by the entity-tag: read with `GET`, write with `If-Match`, and a stale write is refused with `412 Precondition Failed`. What the portal gives up — locks, staged multi-request changes, confirmed commit — it rarely needs. ## The case for NETCONF in the window A NETCONF orchestrator can run, on each device, the sequence RFC 6241 Appendix E describes for multi-device changes: 1. `<lock>` running and the candidate on every target device and keep the locks until the change is permanent or abandoned. 2. Load the edits into each candidate with `<edit-config>` and `<validate>` everywhere. 3. Issue a **confirmed commit** on each device, with `<persist>` if anything else might commit meanwhile. 4. Test the network. 5. Send the confirming commit everywhere — or `<cancel-commit>`, or simply let the timeout revert the trial, 600 seconds unless the orchestrator sets another. 6. `<unlock>`. RESTCONF could push the same edits, but with no lock, no staging across requests and no timed revert, every failure becomes compensation code. ## What neither protocol gives RFC 6241 Appendix E.2 is direct: "Providing complete transactional semantics across multiple devices is prohibitively expensive, but the size and number of windows for failure scenarios can be reduced." Each device commits on its own; no protocol message spans devices. So the orchestrator must plan for partial outcomes: - If 37 of 40 devices are confirmed and the session to the other 3 drops, those 3 revert — a confirmed commit without `<persist>` is undone when its session ends or its timeout expires — and the network is now mixed. - RFC 6241 separates changes that may fail per device and be retried (adding an NTP server) from changes that must land everywhere or nowhere (a VPN or class-of-service definition). Only the second class justifies the full lock-validate-confirm choreography. - Device capabilities vary: check each `<hello>` for `:candidate`, `:confirmed-commit:1.1`, `:validate` and `:rollback-on-error`, and plan a weaker path for devices without them. ## Running both on the same devices If the portal and the window share devices, the choice is no longer independent. On a device without a writable running datastore each RESTCONF edit triggers a commit (RFC 8040 §1.4), which would publish the window's unfinished candidate edits or confirm its trial commit. The window must therefore lock both datastores and persist its confirmed commits; the portal must treat `409 Conflict` as "maintenance in progress". An alternative is a single path: the portal hands its edits to the same orchestrator, at the cost of latency and of losing the simple direct API. ## How to answer in an interview - Name the **unit of change** and the **cost of a partial change** for each workflow. - Give RESTCONF the portal and NETCONF the window, and say what each choice gives up. - State that **cross-device atomicity is the orchestrator's job**, with RFC 6241's techniques narrowing, not closing, the failure windows. - Point out the **coexistence rule** before anyone asks.
- Why not make the portal speak NETCONF as well, for one protocol everywhere?It can, and some teams prefer a single client stack and access model. The costs are session handling, an XML client and SSH or mutual-TLS identities for a web service, for edits that need no lock or staging. Whichever protocol the portal uses, it still collides with the window on shared devices; one protocol simplifies tooling, not coordination.
- How does the window handle devices that lack :candidate or :confirmed-commit:1.1?Read each device's hello and put those devices in a separate risk class. On a running-only device, lock running, save the current configuration with `<get-config>` as a checkpoint, and use `rollback-on-error` if `:rollback-on-error` is advertised. Without a confirmed commit there is no automatic revert, so order the changes so management reachability is never cut first, and keep another path to the device.
saying these in an interview costs you the question
- NETCONF commits a change atomically across every device it touches.
- Confirmed commit is a network-wide rollback mechanism.
- RESTCONF has no transactions at all, so it is unfit for production changes.
- Once each workflow has its own protocol, the two need no coordination.
- Every NETCONF device supports the candidate and confirmed commit.