skip to content

Design a resilient Config Client startup: reconcile fail-fast, retry, optional:, and precedence for a service that must not boot with wrong config but should survive Config Server restarts. What trade-offs do you weigh?

level: principalimportance: nice to knowfreq 34%

answer

  1. non-optional + fail-fast + retry = resilient boot
  2. retry budget < startup/readiness probe timeout
  3. override-system-properties=false lets secrets win
  4. optional: masks outages — avoid for hard deps
  5. runtime refresh is a SEPARATE axis

basics

~20 s

Use a non-optional configserver: import with fail-fast=true plus spring-retry so brief outages are retried but a truly missing server aborts startup. Keep secrets from env vars winning via override-system-properties=false, and let central config override local defaults.

solid answer

~50 s

For a service where wrong config is worse than not starting, I make the import non-optional (configserver:, not optional:) and set fail-fast=true so an unreachable server aborts boot loudly. To ride out rolling restarts I add spring-retry + spring-boot-starter-aop and tune retry (max-attempts, initial-interval, multiplier, max-interval) so total retry time fits inside the orchestrator's startup/readiness probe window — otherwise the pod is killed mid-retry. For precedence, central config stays authoritative (default remote-wins) for shared settings, but I set override-system-properties=false so injected secrets/env vars beat stale central values. I avoid optional: for hard dependencies because it masks outages. I also plan runtime refresh (@RefreshScope, /actuator/refresh or Bus) separately, and monitor with the config-server health indicator. The core tension: fast, loud failure vs. tolerance of transient unavailability — retry backoff bounded by probe timeouts is the lever that balances them.

code

yaml · 19 lines
yaml
spring:
  application:
    name: orders
  config:
    import: "configserver:http://config:8888"   # non-optional: real dependency
  cloud:
    config:
      fail-fast: true                             # abort on unreachable server
      override-system-properties: false           # env/JVM secrets beat central config
      retry:
        max-attempts: 6                           # keep total backoff < startupProbe window
        initial-interval: 1000
        multiplier: 1.1
        max-interval: 2000
management:
  endpoint:
    health:
      show-details: always                        # surface config-server health indicator
# deps: spring-cloud-starter-config, spring-retry, spring-boot-starter-aop

go deeper

for a junior

Recognize that fail-fast plus retry is the resilient combo.

for a middle

Explain why optional: is wrong for a hard dependency and that retry needs its deps.

for a senior

Tune retry backoff and precedence flags for a concrete deployment.

for a principal

Reconcile all levers against orchestrator probe timeouts and secret-injection strategy, and separate startup resilience from runtime refresh in the overall config architecture.

## The competing goals 1. **Correctness over availability at boot:** a service with the wrong DB URL or feature flags is dangerous; failing to start is safer. 2. **Availability across normal churn:** Config Servers get redeployed; clients that start during that window shouldn't die permanently. These pull in opposite directions. The design job is to fail **fast and loud** for real problems while **absorbing transient** ones. ## The building blocks and how they combine ### Non-optional import + fail-fast ``` spring.config.import=configserver:http://config:8888 # NOT optional: spring.cloud.config.fail-fast=true ``` - Non-optional means a load failure is an error (not silently skipped). - `fail-fast=true` makes that error **fatal** to startup. - `optional:` would tolerate a missing server — appropriate only when config is genuinely nice-to-have. For a hard dependency, `optional:` is a footgun that lets the app boot mis-configured. ### Retry to absorb transients Add `spring-retry` + `spring-boot-starter-aop`, then bound the backoff: ``` spring.cloud.config.retry.max-attempts: 6 spring.cloud.config.retry.initial-interval: 1000 spring.cloud.config.retry.multiplier: 1.1 spring.cloud.config.retry.max-interval: 2000 ``` **Critical constraint:** the *total* worst-case retry duration (sum of intervals across attempts) must be **shorter than the platform's startup/readiness probe timeout**. If Kubernetes' `startupProbe` failureThreshold × periodSeconds is 60s but your retries can run 90s, the orchestrator kills the pod before retry succeeds — you get flapping instead of resilience. Compute the retry budget explicitly. Remember retry only engages **with fail-fast=true**. ### Precedence policy - Keep **remote-wins** (default) for shared platform config — central is authoritative. - Set **`spring.cloud.config.override-system-properties=false`** so orchestrator-injected env vars and JVM system properties (e.g., per-pod secrets, region overrides) beat the central file. This lets you patch a bad central value without a Config Server change. - Consider **`override-none=true`** only for teams that want local-as-source-of-truth with central defaults — usually not for platform services. ## Beyond startup - **Runtime refresh** is a separate axis: `@RefreshScope` beans + `POST /actuator/refresh`, or Spring Cloud Bus for fleet-wide refresh. Don't conflate startup retry with runtime refresh — the resilient boot design says nothing about picking up later changes. - **Health & observability:** the Config Server client contributes a health indicator (surfaced at `/actuator/health`); export it so you can alert when clients can't see the server. Log the resolved config coordinates (name/profile/label) at startup for auditability. - **Secrets:** don't rely solely on plaintext central config; combine with a secrets backend (Vault) or env-injected secrets, which is exactly why `override-system-properties=false` matters. ## Failure-mode matrix (design tool) | Scenario | optional: | fail-fast | retry | Result | |---|---|---|---|---| | Hard dependency, want loud failure + churn tolerance | no | true | yes | Retries transient; aborts real outage. **Recommended.** | | Config nice-to-have | yes | false | n/a | Boots with local defaults if server absent. | | Naive hard dependency | no | true | no | Any 1s blip kills startup — fragile. | | Silent misconfig risk | yes | true | — | Contradictory; optional masks the failure fail-fast wants. Avoid. | ## The one-line principle Bound your retry budget by the orchestrator's probe window, make the import non-optional with fail-fast, and let env-injected secrets override central config — that's the sweet spot between *don't boot wrong* and *survive a redeploy*.

  • Why is bounding the retry budget against the orchestrator's startup probe the key constraint?
    If worst-case total retry time exceeds the startup/readiness probe timeout, the platform kills the pod mid-retry, so retries never get to succeed — you get crash-looping instead of resilience. The backoff must fit inside the probe window.
  • Why prefer override-system-properties=false in a Kubernetes deployment?
    So per-pod injected env vars / secrets take precedence over the central config, letting you override or hotfix values without changing the Config Server, and keeping secrets out of shared central files.
  • Does this resilient-startup design address picking up config changes after boot?
    No — that's a separate concern handled by @RefreshScope plus /actuator/refresh or Spring Cloud Bus. Startup retry and runtime refresh are orthogonal.

saying these in an interview costs you the question

  • Using optional:configserver: for a service that genuinely requires central config
  • Setting huge retry budgets that exceed the orchestrator's startup probe, causing crash loops
  • Assuming central config should always beat injected secrets (should set override-system-properties=false)
  • Conflating startup retry with runtime @RefreshScope refresh

context