skip to content

Which configuration and secret differences between staging and production are acceptable?

level: middleimportance: should knowfreq 57%

answer

  1. Two groups of settings, not one
  2. Environment-naming values may differ
  3. Behaviour-changing values should not
  4. A missing key must fail loudly
  5. Compare resolved settings, not files

basics

~20 s

Values that name the environment — endpoints, credentials, account identifiers, resource sizes — are meant to differ. The key set, defaults, precision settings, timeouts and feature-flag states are not. Divergence in behaviour-changing settings is what escapes to release day.

solid answer

~50 s

Split the settings into two groups. Environment-naming values — hostnames and base URLs, connection strings, credentials, external account identifiers, resource sizing — must differ, and that is fine. Behaviour-changing values — numeric precision and rounding mode, currency and locale settings, timeouts, retry and back-off counts, page and batch sizes, cache lifetimes, log levels that switch code paths, and feature-flag state — should match unless somebody decided otherwise and wrote it down. Keep one templated definition with a per-environment value file so both carry an identical key set by construction, make a missing required key fail at startup instead of falling back to a built-in default, and never branch on the environment's name inside application code. Then compare the *resolved* settings of each running instance rather than the source files. Secrets follow the same rule: same names, same injection mechanism, different values.

code

pseudocode · 17 lines
pseudocode
REQUIRED = [
    "billing.rounding.scale",
    "billing.currency",
    "queue.visibility.timeout.seconds"
]

function start(config):
    missing = []
    for key in REQUIRED:
        if not config.has(key):
            missing.append(key)
    if missing.size() > 0:
        abort("missing required settings: " + join(missing, ", "))

    scale = config.integer("billing.rounding.scale")
    log("effective billing.rounding.scale=" + scale)
    return run(config)

go deeper

for a junior

Know that endpoints, credentials and account identifiers are supposed to differ, while timeouts, rounding settings and flag states are not. Be able to give one concrete bug caused by a setting that differed between environments.

for a middle

Explain the mechanics you would put in place: one templated definition with a per-environment value file, required-key validation at startup, effective values logged, and a comparison of resolved settings rather than source files.

for a senior

Show how you would find such a divergence during an incident — dump the effective settings from the running instance, compare them with the other environment, and distinguish a differing value from a key that quietly fell back to a built-in default.

for a principal

Own the policy: which classes of setting may differ at all, who approves an exception, how feature-flag state is kept aligned across environments, and how the rule is enforced automatically rather than by reviewer attention.

## Two groups of settings Configuration divergence is the usual way two environments that look structurally identical stop behaving identically. Start by splitting settings into two groups and treating them under different rules. **Values that name the environment must differ.** Connection strings, hostnames and base URLs, credentials, external account identifiers, storage locations, queue names, resource sizing. These are what make an environment a distinct environment at all. Demanding that they match is meaningless — nobody wants the pre-release environment writing to the production data store. **Values that change behaviour should match**, unless somebody deliberately decided otherwise and recorded it: numeric precision and rounding mode, currency and locale settings, timeouts, retry counts and back-off, page sizes and batch limits, cache lifetimes, log levels that switch code paths, and the state of every feature flag. When one of these differs, the run before release exercised different behaviour from the one that ships — and it still passed, which is exactly what makes this failure mode so quiet. ## The silent default The most common mechanism is a setting that is present in one environment's value file and absent from the other's, where the code carries a built-in fallback. Both environments start cleanly. Nothing warns. The two simply compute different answers. The countermeasures are mechanical. Keep one templated definition with one value file per environment, so the *key set* is identical by construction and only values vary. Validate required keys at startup and refuse to start when one is missing, rather than defaulting. Log the effective value of behaviour-changing settings at startup, so the value actually in force is discoverable during an incident instead of inferred from source. And do not branch on the environment's name inside application code: a branch guarded by "if this is production" is by definition never exercised anywhere else, which places untested behaviour in the one place you cannot afford it. ## Compare resolved settings, not files Effective configuration is normally assembled from several layers: a base definition, a per-environment value file, injected variables, a remote settings store, and built-in defaults. Comparing source files therefore compares the wrong artefact. What you want is each running instance's *resolved* view — dump it with secret values redacted but their keys kept, and compare key by key across environments. Two categories deserve an alert: a key present in one environment only, and a behaviour-changing key whose values differ without a recorded exception. ## Secrets Secrets follow the same rule with a twist: same names, same injection mechanism, same rotation path, different values. The traps are shape and scope rather than value. A pre-release credential may have a different length or prefix, so a validation rule that passes in one environment rejects the other. A credential may carry a narrower permission scope, which quietly turns a secret difference into a permission gap. And a secret delivered as a file in one environment and as a variable in another can pick up a trailing newline in exactly one of them, producing an authentication failure that looks like a wrong password. ## Flag state Feature-flag state is configuration that usually lives outside the configuration system, which is why it is so often left out of this discussion. If the pre-release environment runs with 23 flags enabled and production runs with 9, the behaviour that was verified is not the behaviour that ships. The working discipline: give a flag the same default state everywhere and flip it deliberately; record the flag state a run observed and attach it to that run's results; and treat long-lived flags as debt, because each one doubles the space of behaviours the two environments can occupy. ## A worked example A media platform bills per transcoded minute from its video-transcoding queue. The rounding scale for the per-minute rate came from a setting. The pre-production value file set that scale to four decimal places; the production file, written earlier, omitted the key entirely, so the code's built-in scale of two applied instead. Both environments started green. The 6-hour nightly billing roll-up completed on both. What surfaced was a currency-rounding drift: each job's charge differed by up to a third of a cent, and across 3.7 million monthly jobs the invoice total stopped reconciling with the ledger. There was no defective line of code — there was one missing key and one silent default. A resolved-configuration comparison would have listed that key as present in one environment only, on the day it was introduced. ## What to say in an interview Name the two groups. Give the silent-default mechanism, because it is the one interviewers are listening for. Say you would compare resolved values rather than files, and that the comparison includes flag state and secret shape, not just application settings. Then land the principle: the goal is not identical configuration, it is *explained* configuration — every difference either expected by rule or raised as a defect.

  • How would you catch a setting that exists in one environment and silently defaults in the other?
    Two layers. At startup, validate a declared list of required keys and refuse to start when one is missing, so the fallback never gets a chance. In the pipeline, dump each environment's resolved settings with secret values redacted and compare key by key, reporting any key present in only one environment. The first layer stops it reaching production; the second explains it when it does.
  • Why is a feature flag in a different state between environments worse than a differing timeout?
    A differing timeout changes when a code path gives up; a differing flag changes which code path exists. The behaviour verified before release is then simply not the behaviour that ships. It is also harder to see, because flag state usually lives outside the configuration system, so a configuration comparison misses it unless you deliberately record the flag state each run observed.
  • What goes wrong when application code branches on the environment's name?
    The production branch is never executed anywhere else, so the one behaviour you most need verified is the one nothing exercises. Such branches also multiply, and they hide from a configuration comparison because both environments carry the same settings — the difference lives in code. Push the difference into a value instead, so the same path runs everywhere with different numbers.

saying these in an interview costs you the question

  • Says the two environments should be byte-identical, credentials included
  • Relies on a built-in default when a key is missing
  • Branches on the environment's name inside application code
  • Compares configuration files rather than the resolved values
  • Treats feature-flag state as unrelated to environment parity
  • Assumes a secret differs only in value, never in shape or scope

context