A central config server outage blocks several services that are mid-deploy from starting, because they can't fetch their configuration at bootstrap. How would you architect configuration delivery so that a config-server outage doesn't cascade into a platform-wide outage?
answer
- config server = availability floor for everyone
- cache last-known-good locally
- fail-fast=false + fallback defaults
- HA config server + thundering herd headroom
- surface degraded-mode as an alert
basics
~20 sDon't make every service's ability to even start depend on one central system being up right now. Give each service a fallback: a locally cached copy of its last-known-good config it can start from if the central source is unreachable.
solid answer
~40 sTreat the config server as an eventually-necessary source of truth, not a hard runtime dependency: services should fetch config once and cache it locally, either baked into the deployment image as a last-resort default or persisted to local disk from the last successful fetch, so a config-server outage delays picking up NEW changes but doesn't prevent a service from starting or continuing to run on its last-known-good config. Architecturally this means running the config server itself as a highly available, independently scaled service with its own health checks and redundancy, decoupling the blast radius of a config-server incident from every dependent service's ability to boot, and treating `fail-fast=false` plus local fallback defaults as the default posture rather than the exception.
go deeper
Should recognize that depending on one central service for startup is risky; doesn't need to design the fallback mechanism itself.
Should know fail-fast is the default and can name turning it off as one lever, even if unsure of the full fallback design.
Should design the local-cache/fallback-default architecture plus HA config server, and identify the thundering-herd risk.
Should reason platform-wide about which config classes justify fail-fast vs graceful degradation, and own the SLO and ownership model for the config server as shared infrastructure across all teams.
## Why the outage cascades The failure described happens because, by default, config clients — Spring Cloud Config clients, for example — treat the config server as a **hard synchronous dependency at bootstrap**: 1. the client calls the server before the application context finishes initializing; 2. if the call fails, the client fails fast and refuses to start. That's a deliberate safety choice, better to not start than run with unknown config, but it becomes dangerous the moment it's applied uniformly across dozens of services with no fallback, because the config server's availability then becomes transitively **the availability floor of everything downstream of it**. ## Why fail-fast was the right instinct The original motivation for fail-fast-if-no-config is sound: starting a payment service with silently-default database credentials, or a missing feature flag, is worse than not starting at all. The problem isn't the fail-fast instinct — it's that the architecture gives the service no other legitimate source of last-known-good config to fall back to. The only path to correctness runs through one live network call to one system, so that system's reliability requirement implicitly becomes "must be as available as the union of every service that depends on it," which across 40+ services and an autoscaling event is a much higher bar than the config server was originally built to meet. ## The fix — a fallback tier The fix is to add a **fallback tier**: - each service caches its last successfully fetched config to local disk, or the image ships a baked-in "safe default" set from build time; - on bootstrap, if the config server is unreachable, the service uses that cached or default config rather than refusing to start — trading "always-correct, possibly-unavailable" for "usually-correct, always-available." This is a real trade-off, not a free win: a service that starts from stale cached config during an incident might run with an outdated feature flag or a timeout tuned for yesterday's traffic pattern, and teams need monitoring that specifically flags "started from cached or fallback config" so nobody mistakes degraded-mode operation for a fully healthy deploy. ## Hardening the config server itself Beyond client-side fallback, the config server itself should be run as a **highly available** service: - multiple replicas behind a load balancer; - backed by a resilient git host or a local mirror so an upstream git outage doesn't take down internal config serving; - with its own SLOs, on-call ownership, and capacity headroom for the **thundering-herd** case where an autoscaling event or mass rolling restart causes many pods to request config simultaneously. Some organizations additionally put a read-through cache, such as a CDN or reverse proxy, in front of the config server so repeated identical requests during a herd don't all hit the origin. ## Failure modes The signature failure mode of not doing this is exactly the scenario given: an unrelated config-server blip during a deploy or autoscale event cascades into a broader outage, as multiple otherwise-healthy services fail readiness or liveness checks and get killed and restarted by the orchestrator, which then retries and hits the same still-recovering config server, worsening the herd — a classic **retry storm**. A subtler failure mode after adding fallback caching is **config drift** going unnoticed: a service silently running for days on stale cached config after an incident, with nobody aware it never picked up a since-fixed value, because "fell back to cache" wasn't surfaced as a first-class alert. ## The general principle behind it This mirrors the general control-plane-versus-data-plane resilience principle used broadly in infrastructure design: - DNS resolvers cache last-known answers so a resolver outage doesn't instantly break every client; - Kubernetes' own kubelet keeps running already-scheduled pods even if the API server is briefly unreachable. Applied to Spring Cloud Config specifically, teams commonly set `spring.cloud.config.fail-fast=false` with bounded retry attempts and a local `application.yml` fallback checked into the service's own repo as the known-good defaults, reserving hard fail-fast behavior only for genuinely security-critical values, like a missing encryption key, where running with a stale value would be worse than not running at all.
- Why is fail-fast=false alone not a complete solution — what else has to be true for it to be safe?Turning off fail-fast just means the service proceeds without fresh config, but if there's no local fallback or defaults, it may start with nulls or framework defaults that are wrong for production, such as an in-memory datasource, which can be worse than not starting at all. It has to be paired with a genuine last-known-good local fallback plus monitoring that flags when a service is running in that degraded mode, so it's a deliberate, visible trade-off rather than a silent one.
- What's a 'thundering herd' in this context, and how does it make a config-server incident worse?A thundering herd happens when many clients simultaneously retry or newly request the same resource at once — here, an autoscaling event or mass rolling restart causes dozens or hundreds of pods to all call the config server for their bootstrap config within seconds. If the config server is already struggling, say recovering from a brief outage, this sudden burst of simultaneous load can overwhelm it further, turning a small blip into a prolonged outage as retries pile on retries.
- How would you decide which config values are safe to fall back on locally versus ones that should still fail fast?The line is usually blast radius and staleness tolerance: operational tuning values like timeouts, pool sizes, or non-critical feature flags are generally safe to run stale for a while, making them good fallback candidates, whereas values whose staleness could cause a security or correctness incident — a rotated database credential, an encryption key, or a compliance-driven kill switch — should still fail fast if unavailable, because running with an outdated or missing value there is actively dangerous rather than merely suboptimal.
It's like requiring every store in a chain to call headquarters before unlocking its doors each morning — if HQ's phone line goes down, every store stays closed, even though each store perfectly well remembers yesterday's price list and could open just fine on that.
saying these in an interview costs you the question
- Thinks the fix is simply 'make the config server never go down'
- Doesn't mention local caching or fallback of last-known-good config
- Treats fail-fast=false as sufficient on its own with no fallback defaults
- No awareness of thundering herd risk during autoscaling or mass restarts
- Doesn't distinguish which config values are safe to run stale vs which must fail fast