How do you decide whether a shared Go config loader rejects unknown JSON keys across a fleet?
answer
- one boolean, two opposite consequences
- who gets blocked when it rejects
- cost of rejecting differs by stage
- no warn-only mode, so build one
- fatal is negotiable, visible is not
basics
~20 sDecide per stage, not per fleet. Reject unknown keys where failing is cheap and early — CI validation and process startup — and report rather than reject on hot reload of a serving process, with the unknown keys always named in logs.
solid answer
~60 sThe tradeoff is that `DisallowUnknownFields` catches an operator's typo before it becomes a silent misconfiguration, and at the same price rejects any key a newer writer added that this binary does not know — which turns a staged rollout into a fleet-wide config failure. So I place strictness by blast radius rather than switching it on everywhere. Strict where failing costs nothing and the feedback is immediate: validating the rendered config in CI, and at process startup before the instance takes traffic. Lenient where refusing costs availability: a hot reload in a serving process keeps its current configuration and logs the offending key rather than exiting. Because the option has no warn-only mode, I get the report without the failure by decoding twice — leniently for the value the process uses, strictly to log what it did not recognise — and I count those log lines, so leniency stays observable. Ordering is the other half: the reader learns a key before any writer emits it, and the loader's schema is owned by the team that ships the loader.
code
go · 12 linesvar c Config
if err := json.Unmarshal(b, &c); err != nil {
return Config{}, err
}
dec := json.NewDecoder(bytes.NewReader(b))
dec.DisallowUnknownFields()
var probe Config
if err := dec.Decode(&probe); err != nil {
log.Printf("config drift: %v", err)
}
return c, nilgo deeper
Understand that rejecting unknown config keys is a choice with costs on both sides, and that the same option that catches a typo can also block a deploy.
Be able to name where the option is applied — CI, startup, reload — and explain why the answer can differ at each point rather than being one setting for the whole system.
Argue the placement concretely: fail closed at startup, keep serving on a failed reload, and show the double-decode pattern that reports unknown keys without failing.
Own the posture end to end: who owns the loader's schema, the reader-before-writer rollout order, the time-boxed escape hatch, and the metric that keeps leniency from becoming permanent blindness.
## The decision, stated honestly `json.Decoder.DisallowUnknownFields()` is a single boolean with two opposite consequences, and no configuration flag settles which one you want everywhere: - **With it on**, a key nobody's struct understands is a hard failure with the key named. A misspelled setting fails the deploy instead of quietly doing nothing. - **Also with it on**, a key that a *newer* writer legitimately added is equally fatal to any binary that predates it. During a staged rollout, the fleet is deliberately running two versions at once, and a config template that gained a field for the new one now refuses to load on the old one. The second consequence is the one that gets a platform team overruled. When a rollout stalls because half the fleet rejects the config, whoever is blocked will demand leniency, and they will get it — usually as a permanent global switch flipped in a hurry, which throws away the first consequence too. ## Place the decision by stage, not by fleet The useful reframing is that "strict or lenient" is not one decision. It is four, at four points where the cost of rejecting differs by orders of magnitude: 1. **Authoring / CI.** Validate the rendered configuration against the loader's struct in the pipeline that produces it. Rejecting here costs a red build. This is where strictness should be maximal, and it is the setting most teams under-use. 2. **Deploy gate.** A `validate` subcommand of the binary being deployed, run against the config it is about to receive. Rejecting here costs a blocked rollout — annoying, recoverable, and exactly the signal you want when a typo is real. 3. **Process startup.** Fail closed. The instance has no traffic yet, an error naming the key is unambiguous, and starting on defaults instead is the silent misconfiguration you were trying to prevent. 4. **Hot reload of a serving process.** Do not fail closed. Keep the configuration already in use, log the named key at error level, and increment a counter. A running service that exits because of a config typo has converted a small problem into an outage. That ladder gives you strictness where it is cheap and tolerance where it is expensive, without ever being blind. ## Getting the report without the rejection Go's option has no middle setting — there is no warn-only mode and no per-key allowlist. You build the middle yourself, and it is cheap: decode the same bytes twice, once leniently into the struct the process will use, once strictly with the result discarded except for its error, which names the key. The process still runs; the unknown key is now a log line and a metric. That pattern is what makes leniency defensible: you are not ignoring drift, you are choosing not to fail on it while still counting it. An unknown-key counter that is normally zero and spikes during a rollout is also the single best evidence for the ordering rule below. ## Ownership and ordering Two organisational rules do more than any flag: - **The loader's schema has one owner.** If the config struct lives in a shared library, the team that ships it owns what keys exist, and the answer to "can I add a key?" is a change to that library, not a change to a template. - **Readers learn a key before writers emit it.** Roll out the binary that understands the key first; only then let the tooling that renders the config emit it. This ordering is what makes strictness at startup survivable during a staged rollout, and it is why the escape hatch is rarely needed if the sequence is enforced. ## The escape hatch, made healthy Someone will eventually be blocked at 2am. Decide now what they are allowed to do, and make it the least damaging option: - A documented, per-process lenient mode that still logs and counts every unknown key — not a silent global default change. - Time-boxed and visible: an alert that fires while any instance runs lenient, so the state is temporary by construction rather than by intent. - A written expectation that the follow-up is a fix to the loader or the template, not a permanent relaxation. What you are protecting is the property that made strictness worth having: that an unknown key is always *named somewhere*. Whether it is fatal is negotiable; whether it is visible should not be.
- A rollout is blocked because older instances reject a newly added config key. What do you allow the on-call engineer to do?Switch the affected processes to the documented lenient mode, which still logs and counts every unknown key, and let an alert fire while any instance is in that state. What I do not allow is a silent global default change, because that removes the check everywhere and nobody notices it never came back.
- Why not simply run lenient everywhere and rely on review to catch typos?Because a typo produces no artefact for review to catch — the file is valid, the decode succeeds, the setting is absent. Review catches what it can see. Strict validation in CI and at startup is what makes the mistake visible at all, which is why leniency is only defensible when paired with a strict pass that logs.
- How does deployment ordering change how much strictness you can afford?Rolling the reader out before the writer emits a new key means an old binary never meets a key it does not know, so startup strictness costs nothing. Without that ordering, every additive config change risks a rejected rollout, and the pressure to disable the check permanently becomes irresistible.
- What signal tells you the posture is working rather than just switched on?An unknown-key counter that sits at zero and spikes only during rollouts, together with strict CI validation that occasionally goes red. Zero red builds and zero counted keys usually means the check is not in the path at all, not that the fleet is clean.
saying these in an interview costs you the question
- Turns strict decoding on fleet-wide with no escape hatch
- Makes a serving process exit when a hot reload is rejected
- Treats lenient as the default without logging unknown keys
- Lets writers emit a key before readers understand it
- Relies on code review to catch config typos
- Flips the global default during an incident and never reverts