Design a resilient Config Client startup: reconcile fail-fast, retry, optional:, and precedence for a service that must not boot with wrong config but should survive Config Server restarts. What trade-offs do you weigh?
answer
- non-optional + fail-fast + retry = resilient boot
- retry budget < startup/readiness probe timeout
- override-system-properties=false lets secrets win
- optional: masks outages — avoid for hard deps
- runtime refresh is a SEPARATE axis
basics
~20 sUse a non-optional configserver: import with fail-fast=true plus spring-retry so brief outages are retried but a truly missing server aborts startup. Keep secrets from env vars winning via override-system-properties=false, and let central config override local defaults.
solid answer
~50 sFor a service where wrong config is worse than not starting, I make the import non-optional (configserver:, not optional:) and set fail-fast=true so an unreachable server aborts boot loudly. To ride out rolling restarts I add spring-retry + spring-boot-starter-aop and tune retry (max-attempts, initial-interval, multiplier, max-interval) so total retry time fits inside the orchestrator's startup/readiness probe window — otherwise the pod is killed mid-retry. For precedence, central config stays authoritative (default remote-wins) for shared settings, but I set override-system-properties=false so injected secrets/env vars beat stale central values. I avoid optional: for hard dependencies because it masks outages. I also plan runtime refresh (@RefreshScope, /actuator/refresh or Bus) separately, and monitor with the config-server health indicator. The core tension: fast, loud failure vs. tolerance of transient unavailability — retry backoff bounded by probe timeouts is the lever that balances them.
code
yaml · 19 linesspring:
application:
name: orders
config:
import: "configserver:http://config:8888" # non-optional: real dependency
cloud:
config:
fail-fast: true # abort on unreachable server
override-system-properties: false # env/JVM secrets beat central config
retry:
max-attempts: 6 # keep total backoff < startupProbe window
initial-interval: 1000
multiplier: 1.1
max-interval: 2000
management:
endpoint:
health:
show-details: always # surface config-server health indicator
# deps: spring-cloud-starter-config, spring-retry, spring-boot-starter-aopgo deeper
Recognize that fail-fast plus retry is the resilient combo.
Explain why optional: is wrong for a hard dependency and that retry needs its deps.
Tune retry backoff and precedence flags for a concrete deployment.
Reconcile all levers against orchestrator probe timeouts and secret-injection strategy, and separate startup resilience from runtime refresh in the overall config architecture.
## The competing goals 1. **Correctness over availability at boot:** a service with the wrong DB URL or feature flags is dangerous; failing to start is safer. 2. **Availability across normal churn:** Config Servers get redeployed; clients that start during that window shouldn't die permanently. These pull in opposite directions. The design job is to fail **fast and loud** for real problems while **absorbing transient** ones. ## The building blocks and how they combine ### Non-optional import + fail-fast ``` spring.config.import=configserver:http://config:8888 # NOT optional: spring.cloud.config.fail-fast=true ``` - Non-optional means a load failure is an error (not silently skipped). - `fail-fast=true` makes that error **fatal** to startup. - `optional:` would tolerate a missing server — appropriate only when config is genuinely nice-to-have. For a hard dependency, `optional:` is a footgun that lets the app boot mis-configured. ### Retry to absorb transients Add `spring-retry` + `spring-boot-starter-aop`, then bound the backoff: ``` spring.cloud.config.retry.max-attempts: 6 spring.cloud.config.retry.initial-interval: 1000 spring.cloud.config.retry.multiplier: 1.1 spring.cloud.config.retry.max-interval: 2000 ``` **Critical constraint:** the *total* worst-case retry duration (sum of intervals across attempts) must be **shorter than the platform's startup/readiness probe timeout**. If Kubernetes' `startupProbe` failureThreshold × periodSeconds is 60s but your retries can run 90s, the orchestrator kills the pod before retry succeeds — you get flapping instead of resilience. Compute the retry budget explicitly. Remember retry only engages **with fail-fast=true**. ### Precedence policy - Keep **remote-wins** (default) for shared platform config — central is authoritative. - Set **`spring.cloud.config.override-system-properties=false`** so orchestrator-injected env vars and JVM system properties (e.g., per-pod secrets, region overrides) beat the central file. This lets you patch a bad central value without a Config Server change. - Consider **`override-none=true`** only for teams that want local-as-source-of-truth with central defaults — usually not for platform services. ## Beyond startup - **Runtime refresh** is a separate axis: `@RefreshScope` beans + `POST /actuator/refresh`, or Spring Cloud Bus for fleet-wide refresh. Don't conflate startup retry with runtime refresh — the resilient boot design says nothing about picking up later changes. - **Health & observability:** the Config Server client contributes a health indicator (surfaced at `/actuator/health`); export it so you can alert when clients can't see the server. Log the resolved config coordinates (name/profile/label) at startup for auditability. - **Secrets:** don't rely solely on plaintext central config; combine with a secrets backend (Vault) or env-injected secrets, which is exactly why `override-system-properties=false` matters. ## Failure-mode matrix (design tool) | Scenario | optional: | fail-fast | retry | Result | |---|---|---|---|---| | Hard dependency, want loud failure + churn tolerance | no | true | yes | Retries transient; aborts real outage. **Recommended.** | | Config nice-to-have | yes | false | n/a | Boots with local defaults if server absent. | | Naive hard dependency | no | true | no | Any 1s blip kills startup — fragile. | | Silent misconfig risk | yes | true | — | Contradictory; optional masks the failure fail-fast wants. Avoid. | ## The one-line principle Bound your retry budget by the orchestrator's probe window, make the import non-optional with fail-fast, and let env-injected secrets override central config — that's the sweet spot between *don't boot wrong* and *survive a redeploy*.
- Why is bounding the retry budget against the orchestrator's startup probe the key constraint?If worst-case total retry time exceeds the startup/readiness probe timeout, the platform kills the pod mid-retry, so retries never get to succeed — you get crash-looping instead of resilience. The backoff must fit inside the probe window.
- Why prefer override-system-properties=false in a Kubernetes deployment?So per-pod injected env vars / secrets take precedence over the central config, letting you override or hotfix values without changing the Config Server, and keeping secrets out of shared central files.
- Does this resilient-startup design address picking up config changes after boot?No — that's a separate concern handled by @RefreshScope plus /actuator/refresh or Spring Cloud Bus. Startup retry and runtime refresh are orthogonal.
saying these in an interview costs you the question
- Using optional:configserver: for a service that genuinely requires central config
- Setting huge retry budgets that exceed the orchestrator's startup probe, causing crash loops
- Assuming central config should always beat injected secrets (should set override-system-properties=false)
- Conflating startup retry with runtime @RefreshScope refresh