skip to content

Some external configuration stores support dynamic refresh, where a running service instance picks up a changed value without a redeploy or restart. Describe two different mechanisms services use to detect config changes - one push-based, one poll-based - and explain what can go wrong when a config change is rolled out to a fleet of hundreds of instances.

level: seniorimportance: must knowfreq 62%

answer

  1. poll: interval + ETag/version check
  2. push: watch/blocking-query/subscription
  3. propagation delay vs store load trade-off
  4. fleet-wide inconsistency window
  5. atomic settings need coordinated rollout, not dynamic refresh

basics

~20 s

Poll-based: each instance periodically asks the store 'anything new?'. Push-based: the store actively notifies instances (via a watch or event) when something changes. At fleet scale, the risk is instances applying the change at different times, so the fleet briefly runs inconsistent config.

solid answer

~50 s

Poll-based refresh has each instance periodically call the config store (e.g., every 30-60s) and compare a version or ETag to detect changes, then reload if different - simple and resilient to missed events, but adds latency before changes take effect and creates load proportional to fleet size. Push-based refresh uses a long-lived watch or subscription (e.g., Consul's blocking queries, etcd's watch API, a Kafka topic for config-change events) so the store notifies instances near-instantly when a key changes, reducing propagation delay but requiring the store to track many open connections and handle reconnection/missed-event edge cases. At fleet scale, the main risk is a window of inconsistency: instances receive and apply the change at slightly different times (network jitter, GC pauses, staggered poll intervals), so for some period the fleet is split between old and new config, which is dangerous for changes that must be atomic across all instances (e.g., a schema-related setting) and can cause hard-to-reproduce bugs that look like a race condition.

go deeper

for a junior

Not typically expected to answer this in depth; should at least grasp that 'config can change without a restart' is a thing and that it takes some time to reach everyone.

for a middle

Should describe both poll and push mechanisms at a basic level and know that propagation isn't instant.

for a senior

Should name concrete implementations, explain the propagation-delay/load trade-off precisely, and identify the fleet-wide inconsistency window as the central risk.

for a principal

Should discuss mitigations at the rollout-strategy level (canarying config changes like code deploys) and know when to avoid dynamic refresh entirely in favor of coordinated restarts for atomicity-sensitive settings.

## What dynamic refresh adds Dynamic refresh solves a narrower problem than the base **External Configuration Store** pattern: not just 'load config at startup' but 'pick up a config change while already running, without a restart.' This matters because restarting hundreds of instances to change one timeout value is slow, disruptive to in-flight requests, and turns a config change into a mini-deployment with all the risk that implies. There are two broad mechanisms for detecting that a change has happened. ## Poll-based refresh **Poll-based refresh** has each running instance periodically ask the store 'has anything changed for me?' - typically on a fixed interval (every 15-60 seconds is common) or with a lightweight conditional check using a version number or `ETag` so the instance can cheaply detect 'nothing changed, skip the reload' without re-parsing the full config payload every time. Spring Cloud Config's `/actuator/refresh` endpoint combined with a scheduled poller, or a sidecar that periodically re-reads a mounted Kubernetes `ConfigMap`, are common implementations. Polling is simple to reason about and inherently **self-healing**: if one poll is missed due to a transient network blip, the next poll a minute later will still pick up the current state, so there's no permanent desync from a single dropped message. The cost is twofold: - **propagation delay** - a change can take up to a full poll interval to reach every instance - **load on the config store** that scales linearly with fleet size and poll frequency, since every instance independently hits the store on its own schedule ## Push-based refresh **Push-based refresh** instead has the store actively notify subscribed instances the moment a value changes, using a long-lived connection or subscription mechanism: | Mechanism | Shape of the subscription | |---|---| | Consul's blocking queries | an HTTP long-poll that returns only when the watched key's index changes | | etcd's native watch API | a streaming gRPC subscription | | a message bus like Kafka | config changes are published as events that a lightweight config-sync sidecar consumes and applies locally | Push mechanisms cut propagation delay to near-instant and avoid the load of constant polling, but they trade that for more operational complexity on the store side: - it must track potentially thousands of open watch connections - it must handle client reconnection gracefully after a network blip - a client that misses a watch event while disconnected needs a way to catch up, usually by falling back to a poll or fetching current state on reconnect - it must avoid a thundering-herd reconnect storm if the store itself restarts and every client reconnects simultaneously ## The window of inconsistency The real danger at fleet scale isn't either mechanism individually - it's the window of inconsistency both create. Because instances receive and apply a change at slightly different times (staggered poll schedules, network jitter, GC pauses delaying event processing, or simply the store notifying subscribers in some non-atomic order), there is always some period during a rollout where part of the fleet is running the old value and part is running the new one. For most settings this is harmless - a log-level change taking effect gradually across the fleet is fine. But for settings that must be **atomic across all instances** - a schema-related flag, a protocol version toggle, a value that two collaborating instances both need to agree on simultaneously (like a distributed lock timeout or a consistent-hashing ring parameter) - this **split-brain window** can cause genuinely hard-to-reproduce bugs: - two instances disagreeing about how to interpret the same request - a coordination protocol breaking because half the fleet expects one behavior and half expects another These bugs look like intermittent race conditions and are notoriously hard to diagnose because the 'cause' (an in-flight config rollout) is transient and often already finished by the time someone investigates. ## The standard mitigations The standard mitigations are: 1. Treat non-atomic-safe config changes with the same caution as a **code deploy** - canary the change to a subset of instances first, monitor, then roll out gradually rather than pushing to 100% of the fleet at once. 2. Make the reload operation itself **atomic and side-effect-free within a single instance** - swap a config object reference rather than mutating fields in place, so a request never sees a half-updated config mid-reload. 3. For genuinely fleet-wide-atomic settings, avoid dynamic refresh entirely and require a coordinated restart or a feature-flag-style gate that's explicitly designed for synchronized cutover instead.

  • Why is swapping an entire config object reference safer than mutating individual fields during a reload?
    Reference swap makes the reload atomic from the perspective of any single in-flight request: a request either sees the whole old config object or the whole new one, never a partially-updated mix of old and new field values. Mutating fields in place can let a request read some updated fields and some stale ones mid-reload, producing an inconsistent internal state that never existed as either the old or new configuration.
  • How would you detect that a fleet-wide rollout of a config change has actually completed everywhere, not just been triggered?
    Have each instance emit a metric or log line tagged with the config version it's currently running, and track the distribution of versions across the fleet (e.g., a dashboard panel or a script that queries a `/health` or `/info` endpoint on every instance) until 100% report the new version. Relying on 'we pushed it' as proof of completion misses exactly the propagation-delay and missed-event cases that cause incidents.

Poll-based refresh is like everyone checking a shared bulletin board once an hour; push-based is like a PA system announcement heard the moment it's made. Either way, if the announcement changes the rules of a game already in progress, players on the field for a few seconds will be playing by different rules than players who just heard the update - that gap is where things go wrong.

saying these in an interview costs you the question

  • Assumes dynamic refresh is instantaneous and atomic across the whole fleet
  • Can't name a concrete poll or push mechanism, only describes them abstractly
  • Doesn't recognize that a partial-fleet inconsistency window is unavoidable, not a bug to eliminate entirely
  • Suggests all config values are equally safe to dynamically refresh regardless of whether they need cross-instance agreement
  • No mention of atomic reload (reference swap) within a single instance

context