skip to content

Across many services and environments, how do you decide which configuration layer owns each setting, and stop override layers from sprawling?

level: principalimportance: should knowfreq 42%

answer

  1. placement follows how a value varies
  2. lowest layer that expresses the variation
  3. one source order for every service
  4. layer count versus explainability
  5. generate deployment config, never hand-edit

basics

~20 s

Classify each setting by how it varies - never, per environment, per instance, per run - and by sensitivity, then place it in the lowest layer that can express that variation. Cap layer count and share one source order across services.

solid answer

~50 s

Placement follows variation. A value that is the same everywhere belongs in the artifact as a default, where it is reviewed with the code. A value that varies by environment belongs in a version-controlled per-environment override set. A value that varies per deployment or per instance belongs in deployment-supplied sources such as environment variables. A sensitive value comes from the platform's dedicated source. A value that changes for one run belongs on the command line. The sprawl controls are organisational: **one** documented source order for every service, a hard cap on layer count, per-environment override sets small enough to read in a diff, and boot-time validation so a dead key cannot linger. The tradeoff you are managing is flexibility against explainability - every layer buys a way to change a value without a deploy and costs everybody a place to check.

go deeper

for a junior

Know that where a setting lives is a decision, not an accident: values that never change ship with the code, values that differ per environment live in that environment's settings, and per-instance values come from the deployment.

for a middle

Be able to justify a placement from how the value varies and how sensitive it is, and explain why a value that is identical everywhere does not belong in every environment's override set.

for a senior

Show the operational consequences - drift between hand-edited environments, dead keys nothing reads, effective-configuration diagnostics - and the guardrails that keep per-environment diffs small and reviewable.

for a principal

Own the tradeoff explicitly: layer count buys change speed and costs explainability, and the right balance depends on how fast and safe your deployment path already is. Standardise source order and diagnostics before standardising keys.

Once a platform has dozens of services and several environments, configuration stops being a per-service detail and becomes a design problem with organisational consequences. The question is not "can this value be overridden?" - in a layered model everything can - but "which layer *owns* it, and who therefore owns changing it?" ## Classify before you place Four questions settle almost every placement: 1. **Does it vary at all?** If the value is the same in every environment forever, it is not configuration in the interesting sense; it is a default that belongs with the code. 2. **Does it vary per environment, or per instance?** Per-environment values can live in a reviewable, version-controlled override set. Per-instance values - identity, locality, an assigned port or shard - can only come from whatever launches the process. 3. **Is it sensitive?** Sensitive values come from the source the platform provides for them and are merged like any other layer; the interesting part here is only that the layer exists and where it ranks. 4. **How fast must it change?** A value that must change without a redeploy needs a layer that a deployment can rewrite; a value that may wait for the next release is better off in the artifact, where the change is reviewed. ## A default placement policy | Kind of setting | Home layer | Why there | |---|---|---| | Invariant defaults, safe fallbacks | Code / packaged file | Reviewed with the code, travels with the build | | Per-environment endpoints and sizing | Version-controlled override set for that environment | Visible in a diff, reproducible, still not in the build | | Per-deployment or per-instance values | Deployment-supplied variables | Only the launcher knows them | | Sensitive values | The platform's dedicated source | Kept out of files and out of the build | | One-off operator overrides | Command-line arguments | Scoped to a single run by construction | The policy's value is not that it is uniquely correct - it is that it is **one** policy. A reader who knows it can predict where any setting lives in any service, and a reviewer can challenge a placement without relitigating the model. ## What sprawl actually costs - **Explainability.** With six layers, explaining one effective value means checking six places. With two, it means checking two. Incident time scales with layer count. - **Drift.** Values supplied outside version control differ between environments with nothing to diff. Whatever the deployment layer sets should be generated from a checked-in description, not hand-edited. - **Shadow ownership.** Every layer a team can edit without review is a place a production change can happen without review. That is occasionally desirable and usually not. - **Dead keys.** Override sets accumulate keys for settings the code no longer reads, and nothing removes them because nothing notices. - **Untested variants.** Each environment-specific override set is a configuration that only its own environment exercises; the more of them there are, the more of your fleet is running something nobody tested directly. ## Guardrails worth mandating 1. **One documented source order**, identical across services, so an operator's intuition transfers. 2. **A layer budget.** Name the layers a service may use and require an argument to add one. 3. **Small, reviewable per-environment diffs.** If an environment's override set approaches the size of the base set, the "same artifact everywhere" claim is no longer true. 4. **Typed binding with boot-time validation and strict unknown-key handling**, so a key nothing reads cannot survive quietly in an override set. 5. **Effective configuration with origins, exposed in every environment**, redacted - the single highest-leverage artifact for this whole area. 6. **Generated, not hand-maintained, deployment configuration**, so per-environment values are diffable and drift is detectable. ## The tradeoffs with no single right answer **Flexibility versus explainability.** More layers means more ways to change a value quickly, in an incident, without a deploy - and a correspondingly worse answer to "why is production different?" Teams that have been burned by slow deploys over-index on layers; teams burned by mystery values over-index on locking everything into the artifact. The defensible position depends on how fast your deployment pipeline is: cheap, fast deploys make extra override layers much less valuable. **Central standard versus team autonomy.** A single enforced model makes the fleet legible and makes tooling possible; it also imposes a shape on services whose needs genuinely differ. The workable compromise is usually to standardise the **source order and the diagnostics** - the parts that make cross-service reasoning possible - while leaving each service free to decide which keys it needs. **Static versus dynamically reloadable values.** Reloading removes restarts from the change path and reintroduces the mid-life failure that boot-time validation was designed to eliminate. Reserve it for the handful of values that genuinely need it, validate on reload exactly as at boot, and define what happens when the new values are rejected. The mark of a healthy configuration strategy is not the absence of layers; it is that anyone can answer, quickly and without a discussion, which layer is supposed to own a given setting and which one actually supplied it today.

  • How do you decide whether a value belongs in a version-controlled override set or in deployment-supplied variables?
    Ask whether it varies per environment or per instance. Per-environment values benefit from being in version control - reviewable, diffable, reproducible. Per-instance values, and anything the launcher alone knows, must come from the deployment. Sensitive values go to the platform's dedicated source regardless of which of the two they resemble.
  • What is the strongest argument against adding another override layer?
    Every layer multiplies the work of explaining one effective value and creates another place a production change can happen outside review. The counter-argument is speed of change, so the decision turns on deployment cost: when a deploy takes minutes and is safe, an extra layer buys very little and costs permanent diagnostic overhead.
  • Why should deployment-layer configuration be generated rather than hand-maintained?
    Because hand-edited per-environment values drift with nothing to compare. Generating them from a checked-in description makes environments diffable, makes a missing key visible before deployment, and gives change review a place to happen for values that otherwise bypass it entirely.
  • What signals that a service's per-environment override set has grown too large?
    When it approaches the size of the base settings, or when a reviewer can no longer tell from the diff what behaviour changes. At that point each environment is effectively a distinct configuration nobody tests except in place, and the promise that one artifact behaves predictably everywhere has quietly expired.

saying these in an interview costs you the question

  • Adds a new override layer for every awkward value
  • Lets each service define its own source order and diagnostics
  • Hand-edits per-environment deployment values with no diff or review
  • Keeps environment-specific values in the base packaged settings
  • Leaves keys in override sets after the code stops reading them
  • Makes many values dynamically reloadable without validating on reload