How would you design YANG model discovery and caching for an automation platform managing thousands of NETCONF devices on differing software releases?
answer
- what makes two schemas the same
- identifiers are per server
- notice change without reconnecting
- several sources, one verified store
basics
~20 sTreat each device's schema as its module set: names, revisions, features and deviations. Cache sources once in a shared store, detect changes per device from the hello's module-set-id or content-id, and fetch missing modules from a URL or <get-schema>.
solid answer
~40 sI would define a device's effective schema as its implemented modules with revision, features and deviations, plus per-datastore schemas on NMDA servers, and fingerprint that myself. `content-id` and `module-set-id` are implementation-specific and only meaningful per server, so they say when to recheck a device, not that two devices match. Sources go into one shared store keyed by name and revision, filled from a curated repository first, then the library's `location` URLs, then `<get-schema>`, and verified against the advertised revision. Change detection compares the hello identifier each session and, for long-lived sessions, listens for `yang-library-update` or `netconf-capability-change`. The tradeoffs are eager versus lazy fetching, trusting device-served modules versus a curated set, and how much device-specific deviation the platform is willing to model rather than reject.
go deeper
Recall that a device's model is more than module names: revisions, features and deviations all change what the device accepts.
Explain how module-set-id and content-id tell a client when to reread the YANG library, and where module sources can be fetched from.
Show a cache keyed on module set, revision checks on fetched sources, and change detection through hello identifiers and library notifications.
Own the tradeoffs: eager versus lazy fetching, device-served versus curated modules, a deviation policy that decides automatability, and access control on discovery data.
## The question behind the question This is a design judgment with no single answer. A good one starts by defining **what a schema is per device**, because everything else (caching, change detection, compatibility) keys off that definition. ## What "the same schema" means Two devices running a module of the same name and revision can still differ: - **Features**: each device lists the optional features it supports, and `if-feature` nodes exist only where the feature does. - **Deviations**: each device can name deviation modules that remove nodes, change types or mark them unsupported. - **Datastores**: under NMDA (RFC 8342 and RFC 8525), each datastore has its own schema; the operational state datastore may implement modules the configuration datastores do not. - **Import-only modules**: revisions used only for type definitions, which affect validation. So the platform's key is a **fingerprint it computes itself** over that whole set. RFC 8525 states that `content-id` is implementation-specific and that the same information need not produce the same value, so it cannot be compared across devices; `module-set-id` from the older RFC 7895 tree has the same per-server scope. Both are excellent **per-device change signals** and useless as **cross-device identities**. ## Detecting change 1. **At session start**: read `module-set-id` (capability `yang-library:1.0`) or `content-id` (capability `yang-library:1.1`, RFC 8526) from the hello. Unchanged means the cached fingerprint still holds; changed means reread the library. 2. **During long-lived sessions**: subscribe to notifications. RFC 8525 defines `yang-library-update` carrying the new `content-id`; a server that also implements RFC 8526, RFC 5277 notifications and RFC 6470's `netconf-capability-change` emits that notification too, because the capability value changes with it. 3. **At software upgrade**: the orchestration that upgraded the device should force a recheck rather than wait for the next session. ## Acquiring sources | Source | Strength | Weakness | |---|---|---| | Curated internal repository | reviewed, shared, fast | must be kept current per release | | Library `location` URL | no load on the session | present only when the server knows one | | `<get-schema>` (RFC 6022) | always matches the device | needs `ietf-netconf-monitoring`; loads the device | A sensible order is repository, then URL, then `<get-schema>`, always with an explicit **version** (an omitted version fails with `data-not-unique` when several revisions are listed). Every fetched module is checked: its newest `revision` must equal the advertised one, and a device-served copy that differs from the repository's copy of the same name and revision is flagged, not silently overwritten. ## Tradeoffs a lead owns - **Eager or lazy.** Fetching every module at onboarding gives fast first use and a heavy onboarding step; fetching on first need spreads load but lets the first change on a new release fail late. - **Device truth or curated truth.** Device-served modules always match the device, but they can carry deviations that quietly break the platform's assumptions. A curated set is predictable but lags new releases. Many platforms accept device truth for reading and require curated compatibility before writing. - **Deviation policy.** Decide which deviations the platform models and which make a device unsupported for automation, rather than discovering that at change time. - **Exposure.** RFC 8525 warns that the library reveals device type and implementation details useful to an attacker. Grant read access to it, and to the schema store, as carefully as to configuration. ## An onboarding flow that follows from this 1. Open a session and record the hello: base capabilities, `:xpath`, the library capability and its identifier, and any RFC 6020 module URIs. 2. Read the YANG library and compute the fingerprint over modules, revisions, features, deviations and datastore schemas. 3. If the fingerprint is already known, link the device to it and stop. 4. Otherwise resolve each missing source in the repository, URL, `<get-schema>` order, verify revisions, and run the compatibility check against the platform's templates. 5. Store the per-device identifier so the next session can skip steps 2 to 4. ## Failure modes to design out - Keying the cache on hardware model or release name, then breaking when a feature set differs. - Comparing identifiers across devices and sharing a schema that does not apply. - Polling the library on every session instead of comparing the identifier. - Trusting a fetched module without checking its revision against the advertisement.
- Why not deduplicate device schemas across the fleet by their YANG library content-id?RFC 8525 defines `content-id` as an implementation-specific identifier that MUST change when the library changes, with no requirement that the same information always yields the same value. Two identical devices can therefore show different values, and nothing makes a value unique across servers. Its job is per-server change detection; cross-device identity needs a fingerprint the platform computes over modules, revisions, features and deviations.
- When would you refuse to automate a device even though its YANG modules were fetched successfully?When its deviations remove or change nodes the platform's change logic depends on, for example a deviation marking a list unsupported or narrowing a type the templates write. Fetching proves the schema is known, not that it is compatible. A platform should evaluate deviations and missing features against what it intends to write, and mark such a device read-only until the gap is handled.
saying these in an interview costs you the question
- Cache schemas per hardware model, since identical hardware runs identical modules.
- YANG library content-id values can be compared across devices to share schemas.
- Same module name and revision means the same schema, whatever the features and deviations.
- A device's module set never needs rechecking once it has been onboarded.
- Every device serves <get-schema>, so no internal module repository is needed.