If the feature-flag evaluation service itself becomes unreachable or slow, what should a client SDK do, and how does the answer differ between a flag guarding a nice-to-have feature versus a flag acting as a kill switch for a risky dependency?
answer
- cache locally, never block hot path on network
- poll vs stream affects staleness window
- default value per call site matters
- kill switch must not depend on the thing it protects against
- thundering herd on reconnect
basics
~20 sIf the flag system can't be reached, the app should fall back to a safe default instead of crashing or hanging, usually the last value it knew, cached locally. For most features, safe means off; for a kill switch protecting against a broken dependency, safe might mean staying off (fail closed) so the risky code never runs by accident.
solid answer
~50 sFlag SDKs are designed so the flag check never becomes a new single point of failure: they fetch and stream flag state in the background and cache it locally, so a runtime flag evaluation is a cheap local lookup, not a network call in the hot path. If connectivity to the flag service is lost, well-built SDKs fall back to the last cached value rather than blocking or throwing, and only if there's no cache yet (cold start) do they fall back to a hardcoded default shipped in code. Whether that default is on or off is a deliberate per-flag decision: fail-closed (default off) for anything risky or new, so an outage never accidentally exposes half-built functionality; and for a kill switch protecting against a bad dependency, the guarded safe fallback path needs to be the one that doesn't depend on live flag-service connectivity at all, so losing that connection doesn't ironically disable the safety valve.
go deeper
Should intuitively grasp that the app shouldn't crash or hang if it can't reach the flag system, and that it should fall back to something safe.
Should describe local caching plus background sync as the mechanism, and know that fresh cold starts use a hardcoded default.
Should articulate the fail-open/fail-closed distinction for a kill switch specifically, and identify at least one concrete failure mode like stale hardcoded defaults or correlated outages.
Should discuss propagation-latency SLAs, monitoring and game-day validation of kill switches, and system-level considerations like thundering-herd reconnects or relay-proxy architectures at fleet scale.
## How a production-grade SDK is architected Mechanistically, a production-grade flag SDK is architected so the actual "should the flag do X" check inside application code is always a cheap, local, synchronous operation -- a lookup into an in-memory (sometimes also on-disk) cache -- never a network round trip on the request's hot path. Underneath that cache, the SDK maintains its own connection to the flag service in the background: - **Polling** on an interval, such as every 30 seconds. - **Holding a persistent streaming connection** (Server-Sent Events or a WebSocket) that pushes updates as soon as someone changes a flag in the dashboard. When application code calls something like `flags.isEnabled` with a `defaultValue` parameter, the SDK returns whatever's in its local cache; if it has never successfully synced yet -- a true cold start, right after process boot before the first fetch completes -- it returns the literal `defaultValue` the caller hardcoded at the call site, because there's nothing else to return. ## Why the flag service must not become a single point of failure This architecture exists to prevent a subtle but serious problem: pulling a piece of release/ops config out of the binary and into an external service must not turn that external service into a new single point of failure for every request in the system. If evaluating a flag meant blocking a request on a synchronous call to the flag service, an outage or even elevated latency in that one dependency would degrade or take down every service that checks flags, which in a flag-heavy codebase is most of them. Decoupling evaluation from the network call, via cache-then-background-refresh, means the flag service can be fully down for minutes and application requests keep flowing using the last known-good state, completely unaffected in the hot path. ## The trade-off: staleness That design choice creates a real trade-off: staleness. Between the moment someone flips a flag in the dashboard and the moment every running instance's local cache reflects that change, there's a window: | Sync mode | Staleness window | |---|---| | Poll-based SDKs | The polling interval | | Streaming SDKs | Typically sub-second | The window is never provably instant across a fleet of instances that might be mid-restart or behind a flaky network segment. For an ordinary release flag this staleness is harmless; for a kill switch being flipped during a live incident, that window is the response latency of your emergency lever, so teams running kill-switch flags at scale often specifically choose streaming SDKs and monitor SDK-to-service connectivity as its own health signal, because a kill switch that takes two minutes to propagate during an incident has failed at its one job. ## Failure modes The failure modes here are concrete and recur across real outages. 1. **Getting fail-open and fail-closed backwards** -- the first and most consequential. Every flag evaluation call typically carries an explicit default: the value returned when there's no cached state and no connectivity. For a normal, not-yet-launched feature, that default should be false or off, so any disruption to the flag service never accidentally exposes half-built functionality. But for a kill switch -- a flag whose on state means disable this risky subsystem -- the code's own default needs careful scrutiny: if the flag service goes down and the SDK falls back to a hardcoded off default meaning kill switch not engaged, that's fine if the switch was off anyway, but if the whole reason the flag service is unreachable is correlated with the same incident that would have required tripping the kill switch, the team is left with no working lever exactly when they need one. The safer design makes the guarded, safe fallback code path require no live flag connectivity at all, rather than depending on the flag system to actively deliver an off signal during the very outage it's meant to protect against. 2. **Cold-start default drift**, a quiet, slow-burning failure: engineers hardcode a true default at a call site once a feature has been at 100% rollout for months and everyone's forgotten the flag is still being checked; then a genuine flag-service outage or a fresh instance boot before first sync causes that code to briefly serve the wrong, stale default instead of the last real production value, producing behavior nobody intended that's hard to reproduce because it only shows up during connectivity gaps. 3. **Third, a thundering-herd reconnect storm**, which self-hosted or gateway-fronted flag services can suffer -- hundreds of instances losing their streaming connection simultaneously, for example during a flag-service deploy, and all reconnecting or re-polling at once, which SDKs mitigate with jittered backoff. ## Where it shows up in the wild A concrete real-world instance of this design: LaunchDarkly's SDKs require callers to pass an explicit default value on every evaluation call, specifically documented as the value used when the SDK cannot reach the service or hasn't initialized yet, and the company additionally offers a Relay Proxy specifically so a local, low-latency, cacheable layer sits between application instances and the flag backend, reducing both blast radius and staleness during backend disruptions.
- Why is it risky to hardcode defaultValue=true at a flag-check call site for a feature that's been fully rolled out and stable for a year?It quietly encodes an assumption that the flag will always resolve to true, but if the flag service or SDK ever hits a cold start with no cache yet, that hardcoded default is exactly what gets served, so if the real intended state has since changed the code silently reverts to the wrong behavior during any connectivity gap. It's a sign the flag has become stale and should probably be removed rather than trusted as a live check.
- What operational signal should a team monitor specifically to catch kill-switch propagation delay before an incident, not during one?SDK-to-flag-service connectivity and heartbeat health plus the last-successful-sync timestamp per instance, ideally alerting if any instance falls behind its polling or streaming SLA. Testing the kill switch's actual end-to-end propagation time periodically, as a game-day or chaos exercise, is the only way to know the real latency rather than assuming the documented SDK behavior holds under production network conditions.
- Why might a team choose polling over streaming for flag evaluation despite the worse staleness window?Polling is operationally simpler -- no persistent connections to manage, easier to reason about through corporate proxies and firewalls, and avoids the thundering-herd reconnect problem streaming connections can create at large fleet scale. Teams accept a coarser staleness window, such as 30-60 seconds, as an acceptable trade-off when their flags aren't being used as low-latency kill switches.
Like a building's fire alarm pull station wired with its own battery backup instead of relying on the building's main power -- the whole point of an emergency lever is that it still works during the very outage (power loss) that might also be the emergency.
saying these in an interview costs you the question
- Assumes flag evaluation is always a live network call on the request path
- Doesn't distinguish the safe default for a normal feature flag versus a kill switch
- Thinks a kill switch is safe even if it depends on the same infrastructure it's meant to protect against
- Unaware that flag SDKs cache locally specifically to avoid becoming a new single point of failure
- No answer for what happens on a true cold start with no cache yet