skip to content

You run 30 Spring Boot services on Kubernetes and want config changes to take effect safely. How do you design ConfigMap/Secret reload, and what are the failure modes?

level: principalimportance: should knowfreq 30%

answer

  1. event+refresh default; restart_context/shutdown for infra beans
  2. least-privilege RBAC, scope to own ConfigMap/Secret
  3. no staging gate -> GitOps validation + canary
  4. env-injected secrets never hot-update
  5. restart strategies need replicas + probes + PDB

basics

~20 s

Enable reload with event+refresh for fast, zero-downtime changes to refresh-safe values, scope watches to each app's own ConfigMap/Secret with least-privilege RBAC, and use restart_context/shutdown (with replicas + probes) for structural changes. Guard against instant bad config with validation and rollout controls.

solid answer

~40 s

Default to `spring.cloud.kubernetes.reload.mode=event` + `strategy=refresh`: watch each service's own ConfigMap/Secret and rebind `@RefreshScope`/`@ConfigurationProperties` live — sub-second, no downtime, ideal for flags, thresholds, log levels. Grant least-privilege RBAC (`get`/`list`/`watch` on only the relevant objects, per-namespace `Role`, not cluster-wide). For values that infra beans capture once (datasource, listeners), route them through refresh-scoped holders or fall back to `restart_context`/`shutdown`, always with multiple replicas + readiness probes + a rolling Deployment so a restart isn't an outage. Key failure modes: instant reload has no staging gate, so a bad ConfigMap edit hits every pod at once — mitigate with schema/GitOps validation, canary via separate ConfigMaps, and PodDisruptionBudgets. Watch drops, missing RBAC (silent no-op), env-injected secrets that never update, and non-idempotent re-init are the other traps. Consider whether Config Server or Vault fits better for some config classes.

code

yaml · 23 lines
yaml
# Least-privilege RBAC: watch only this app's ConfigMap + Secret
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  namespace: prod
  name: orders-config-reader
rules:
  - apiGroups: [""]
    resources: ["configmaps", "secrets"]
    resourceNames: ["orders-service"]   # scope down to the exact objects
    verbs: ["get", "list", "watch"]      # watch needed for event mode
---
# App: fast path for flags, restart for structural config
spring:
  cloud:
    kubernetes:
      reload:
        enabled: true
        mode: event
        strategy: refresh   # switch to restart_context for datasource-class changes
      secrets:
        enableApi: false    # prefer mounted secrets over API reads
        paths: ["/etc/secrets"]

go deeper

for a junior

Know reload can be instant (event) or restart-based, and that changes should be reviewed.

for a middle

Pick event+refresh for flags and restart strategies for infra beans; know RBAC is required.

for a senior

Design least-privilege RBAC, scope watches, and identify non-refreshable beans and env-secret limits.

for a principal

Own blast radius: no-staging-gate mitigation via GitOps/canary, thundering-herd, watch-drop semantics, PDBs, and per-config-class choice of ConfigMap vs Vault/Config Server.

**Framing.** Config reload without a Config Server is attractive operationally — Kubernetes is already the source of truth — but at fleet scale the trade-offs are about blast radius, RBAC, and which changes are truly hot-swappable. **1. Default posture: `event` + `refresh`.** Watch-based detection gives near-instant propagation; `refresh` rebinds `@ConfigurationProperties` and recreates `@RefreshScope` beans with zero restart. This suits the majority of config: log levels, feature flags, timeouts, rate limits, thresholds. It's the same `ContextRefresher` machinery as `/actuator/refresh`. **2. Least-privilege RBAC.** The API-based config/secret access and watch require the pod's service account to have verbs on `configmaps`/`secrets`. Scope a namespaced `Role` to the specific object names where possible (resourceNames), grant only `get`/`list`/`watch`, and bind per-service accounts — never a cluster-wide, all-secrets grant. For Secrets, prefer mounted volumes (`secrets.paths`, `enableApi=false`) so you don't hand the app broad secret-read at all; only enable the API path when you truly need dynamic secret reload. **3. Scope the watch.** Monitor each app's own ConfigMap/Secret rather than everything in the namespace, to cut noise and avoid unrelated churn triggering refreshes. Toggle `monitoring-config-maps`/`monitoring-secrets` deliberately. **4. Non-refreshable config → heavier strategies.** Datasources, connection pools, Kafka listeners, thread pools often build once and ignore refresh. Options: (a) wrap the holder in `@RefreshScope` and rebuild the resource on recreation (careful with connection churn), (b) `strategy=restart_context` to reboot the Spring context, or (c) `strategy=shutdown` and let the Deployment restart the pod. Both (b)/(c) require multiple replicas, readiness/liveness probes, a rolling update strategy, and ideally a `PodDisruptionBudget` so reload doesn't drop below quorum. **5. Failure modes and mitigations:** - **No staging gate.** An `event`-mode edit hits all pods instantly with no review. Mitigate with GitOps (Argo/Flux) so ConfigMaps are PR-reviewed + schema-validated before apply; roll out via canary ConfigMaps or progressive delivery; keep the change auditable. - **Silent no-op.** Missing `watch` RBAC means reload just doesn't fire — no loud error. Add health/log checks; test with `/actuator/refresh` locally to confirm beans are refresh-aware. - **Watch reconnection gaps.** Watches drop; the client re-lists on reconnect, but design for eventual, not guaranteed-instant, delivery. - **Env-injected secrets never update** in a running container regardless of strategy — only property-source values and volume mounts change. Don't promise dynamic rotation for env-var secrets. - **Non-idempotent re-init.** Refresh re-runs bean init; side-effecting `@PostConstruct` can double-execute. Keep it idempotent and cheap. - **Precedence surprises.** Same key in a Secret and ConfigMap — refresh preserves source ordering; document winners. - **Thundering herd.** Fleet-wide simultaneous refresh can spike downstream (e.g. all pods rebuild a client at once). Stagger or use polling with jitter for large fleets. **6. When NOT to use direct ConfigMap reload.** If you need versioned config history, audit, encryption, or cross-cluster config, a Config Server or Vault may serve some config classes better; you can mix — Kubernetes ConfigMaps for env-shaped flags, Vault for secrets. Choose per config class, not one-size-fits-all. **7. Testing/validation.** Validate refresh-awareness with `POST /actuator/refresh` in staging; run a game-day pushing a bad ConfigMap to confirm blast radius and rollback path (revert the ConfigMap → another RefreshEvent restores prior values, since Kubernetes is the source of truth). **Bottom line.** `event`+`refresh` + least-privilege RBAC + refresh-safe beans for the common case; `restart_context`/`shutdown` behind replicas+probes for structural config; GitOps validation and canarying to compensate for the missing staging gate.

  • A ConfigMap edit with a bad value just rolled out to all 30 pods instantly. How do you prevent and recover?
    Recover by reverting the ConfigMap — Kubernetes is the source of truth, so a revert fires another RefreshEvent restoring prior values (or roll the Deployment for restart strategies). Prevent it with GitOps: ConfigMaps as PR-reviewed, schema-validated manifests; canary via a separate ConfigMap/namespace; progressive delivery; PodDisruptionBudgets so restart strategies never drop below quorum.
  • When would you choose `restart_context` or `shutdown` over `refresh` despite the downtime?
    When the affected beans can't rebind cleanly — datasources, connection pools, message listeners, thread pools built once at context init. Pair with multiple replicas, readiness probes, and rolling updates so the restart is invisible to callers.

saying these in an interview costs you the question

  • Granting cluster-wide secret read to enable reload
  • Promising dynamic rotation for env-var-injected secrets
  • Assuming instant reload is safe without any validation/canary
  • Using restart_context/shutdown without replicas and probes

context