Across a fleet of Go services wired with flags and env vars, which source wins, and when must startup fail?
answer
- specificity decides the order
- consistency beats the choice itself
- classify settings, do not pick one posture
- what is unsafe to guess must abort
- fail-fast risks a fleet-wide crash loop
basics
~20 sPick one order for every service, usually flag over environment over built-in default, and implement it once. Then classify settings: anything unsafe to guess must abort startup with a nonzero exit, while settings with a documented safe default may boot and log what they used.
solid answer
~50 sTwo decisions, and consistency matters more than either answer. First, the order: I standardise on flag beats environment beats compiled-in default, because the flag is the most deliberate thing a human typed and the environment is the platform's channel. What matters is that every service resolves identically and prints the effective values at startup, so nobody has to read the code to know which layer won. Second, fail-fast: I classify each setting. Anything unsafe to guess — a database address, a signing key, a region — must abort with a nonzero exit before the process listens. Anything with a documented safe default — log level, cache size — may boot and log its value. The real cost is that fail-fast turns one bad value into a fleet-wide crash loop, so I pair it with validating config in the pipeline and a rollout that halts on the first failing instance. That trade is the part a platform team can legitimately overrule.
code
go · 9 linesfunc main() {
cfg, err := loadConfig(os.Args[1:]) // defaults, then environment, then flags
if err != nil {
fmt.Fprintln(os.Stderr, "config:", err) // every problem, not just the first
os.Exit(1) // before we listen: a bad deploy never serves
}
logEffective(cfg) // secrets redacted
log.Fatal(http.ListenAndServe(cfg.Addr, newServer(cfg)))
}go deeper
Be able to say which order the service you worked on used and where that is written down. Knowing that a flag normally beats an environment variable, which beats a compiled-in default, is enough here.
Explain how the order is implemented in Go — seeding flag defaults before parsing — and be able to argue why a missing datastore address should stop the process while a missing log level should not.
Show the operational half: validation that runs before the process listens, one complete error report rather than the first failure, and an effective-config line at startup with secrets redacted.
Own the fleet-wide call and its cost. Justify one order for every service, defend which settings abort startup, and explain how you keep fail-fast from converting a single typo into a fleet-wide crash loop.
## What is actually being decided Go gives you no configuration framework, so every service invents its own layering out of `flag`, `os.LookupEnv` and a few `strconv` calls. Left alone, twenty services grow twenty resolution orders and twenty opinions about what a missing value means. The decision is not really technical — any of the plausible orders works — it is about making one choice, implementing it once, and making the outcome visible. ## The precedence order The order almost everyone converges on is **command-line flag, then environment variable, then compiled-in default**. The reasoning is specificity: a flag is something a human deliberately typed for this one invocation; the environment is what the platform injected for this deployment; the compiled default is the fallback that lets `go run ./cmd/server` work on a laptop. In Go this order costs nothing to implement, because a flag's default is an ordinary argument. Resolve the environment first, pass the result as the default, and let `flag.Parse` overwrite it only for flags that actually appeared. When a layer must be applied *after* parsing — a config file whose path is itself a flag — `flag.Visit` tells you which flags were explicitly set so the later layer can skip them. What you should refuse is per-service exceptions. The moment one service inverts the order "because our operators prefer it", every incident involving that service costs an extra ten minutes. ## Fail fast or boot on a default The second decision is sharper, and it is genuinely a trade. **Fail fast** means the process validates every resolved setting and, on any problem, prints to stderr and exits nonzero before it listens, dials a database, or registers with service discovery. The gain is that a misconfigured instance never serves a request: no half-configured process, no requests answered with the wrong feature flags, no data written to the wrong place. The failure is loud, immediate, and attributable to the deploy that caused it. **Boot on a default** means the service starts with a fallback and logs a warning. The gain is availability: a partially broken configuration does not take the service down. The way to hold both is to classify, per setting, rather than to pick one posture for everything: - **Must abort.** Anything where a guess is wrong in a way that is invisible or dangerous: the address of a datastore, a credential, a region or tenant identifier, and anything that changes what the service *means*. There is no safe default for "which database", and a fallback to localhost is how a service quietly writes production data nowhere. - **Must abort even though it parses.** Values that are syntactically fine and semantically impossible: a non-positive request timeout, a rollout percentage above 100, a worker count of zero. Range validation belongs beside type validation. - **May default.** Settings that affect only degree, where the default is documented and logged: log verbosity, cache size, batch size, the port in a development build. ## The organisational cost of fail-fast, and how it is paid The honest objection to fail-fast is real: a config change is usually applied to the whole fleet at once, so a fail-fast service converts one typo into every instance crash-looping, and a crash-looping fleet cannot be rolled back by the same system that just broke it. This is where the decision leaves engineering taste and becomes policy. The mitigations are all outside the process: - **Validate before deploying.** The same resolution and validation code should be runnable as a check in the pipeline against the environment about to be applied. If validation lives in a function that takes an argument slice and an environment lookup rather than reading globals, this is nearly free. - **Roll out progressively and halt on the first failure.** One instance failing to start must stop the rollout rather than proceed to the rest. - **Keep the last-good configuration reachable** so the rollback path does not depend on the thing that broke. - **Make the error report complete.** Collect all validation failures with `errors.Join` and print them together, naming each setting and its offending value. Failing on the first problem forces the engineer bringing up a new environment through one deploy per mistake. A platform team may reasonably overrule a service owner here — for instance requiring that a specific setting have a safe default because the fleet cannot tolerate a simultaneous restart on it. That is a legitimate call, and the service owner's job is to make the consequence explicit rather than to litigate the default. ## Making it stick Ship the policy as a small internal package rather than a document: one function that resolves the layers, one that validates, and a table-driven test over the precedence cases so a future refactor cannot silently reorder them. Require every service to log its effective configuration once at startup with secrets redacted — that line is what turns "it behaves differently in staging" into a ten-second diagnosis. And keep `os.Exit` in `main`. It does not run deferred functions, so a helper that exits on its own skips every cleanup its callers registered and is untestable besides. Helpers return errors; `main` decides to die.
- A platform team insists one setting must always have a fallback rather than aborting. How do you respond?Accept it if the reasoning is about restart blast radius, and make the consequence explicit: document the fallback, log loudly whenever it is used, and add an alert on that log line so a silent fallback does not become permanent. What I would resist is a blanket rule, because a credential or a datastore address has no safe fallback at all.
- How do you stop fail-fast from turning one bad value into a fleet-wide outage?Move the check earlier and slow the blast radius. Run the same resolution and validation code as a pipeline check against the configuration about to be applied, roll out progressively and halt on the first instance that fails to start, and keep the last-good configuration reachable so rollback does not depend on the broken path.
- What do you require every service to emit at startup, and why?One line per setting with the effective value and secrets redacted, printed before the service listens. It is the fastest answer to 'why is this instance behaving differently', it makes a zero or empty value obvious immediately, and it removes the need to read resolution code during an incident.
- Why implement this as a shared internal package rather than a written convention?A convention drifts the moment somebody is in a hurry; a package with a table-driven test over the precedence cases cannot silently reorder its layers. It also gives you one place to add a policy change — a new redaction rule, an extra validation — instead of twenty pull requests across twenty services.
saying these in an interview costs you the question
- Lets each service pick its own precedence order
- Applies one posture to every setting instead of classifying them
- Falls back to a localhost datastore address when the real one is missing
- Fails on the first invalid setting instead of reporting them all
- Adopts fail-fast without a progressive rollout or a pipeline check