skip to content

A globally distributed service uses DNS-based health checks, such as Route 53 health checks, to automatically fail traffic away from an unhealthy region. What makes this pattern riskier than an in-cluster readiness probe, and how would you design against flapping and split-brain?

level: principalimportance: nice to knowfreq 30%

answer

  1. DNS caching layers outside your control
  2. TTL is a request, not a guarantee
  3. asymmetric hysteresis: fast fail-over, slow fail-back
  4. failback thundering herd
  5. regional failover has bigger blast radius, slower action than pod-level

basics

~20 s

DNS-based failover switches which region's address gets handed out based on health checks, but DNS answers get cached everywhere, in browsers, ISPs, and apps, for minutes, so recovery is slow and uneven. If the health check itself flaps, users can end up split across two different regions at once, which is riskier than a fast, uncached, in-cluster check.

solid answer

~60 s

DNS-based health checks work by having the DNS provider poll an endpoint in each region and, on failure, stop returning that region's address in DNS answers, shifting new lookups to a healthy region. This differs fundamentally from a Kubernetes readiness probe in propagation and caching: DNS answers are cached at layers outside the provider's control, client OS resolvers, ISP resolvers, corporate proxies, often ignoring or exceeding the configured TTL, so a failover or failback can take minutes to reach all clients, and different users can be split across regions simultaneously during that window. It's riskier because both the blast radius and the latency of the decision are much larger: a false positive doesn't just pull one pod out of rotation for a few seconds, it can strand a fraction of global traffic on a stale answer for the TTL window. Designing against this means asymmetric hysteresis, requiring sustained failure to fail over but a much longer sustained recovery before failing back, and treating regional failover as a deliberately slower, more conservative action than pod-level remediation.

go deeper

for a junior

Not expected to know this pattern in depth; credit for recognizing that DNS answers get cached and that matters for how fast failover works.

for a middle

Should know that DNS TTL is only a hint to clients and real propagation can lag it substantially.

for a senior

Should articulate the caching-layer mechanics and propose basic hysteresis, not failing back instantly, as a mitigation.

for a principal

Should design the full asymmetric hysteresis policy, reason about failback thundering herds and stateful split-brain, and weigh DNS-based failover against alternatives like global anycast load balancers with their respective cost and propagation trade-offs.

## How the mechanism works, and what happens after the decision A DNS-based health check works by having the DNS provider, such as Route 53, poll a regional endpoint on a fixed interval and mark it healthy or unhealthy; a failover, weighted, or latency-based routing policy then removes the unhealthy region's records from the answers it serves to new lookups. The critical difference from an in-cluster mechanism is what happens after that decision is made. A DNS answer carries a TTL, but every resolver in the actual request path: - the client's OS stub resolver; - a local caching resolver; - an ISP's resolver; - a corporate proxy; - and sometimes an application that reuses a connection rather than re-resolving per request, can hold onto that answer well past its stated TTL, since many of these layers don't honor very short TTLs strictly or impose their own minimum caching duration. ## Why anyone uses it DNS-based failover exists as the traditional lowest-common-denominator mechanism for steering traffic across independent regions or even clouds, where there's no single load balancer spanning them the way a Kubernetes Service spans pods within one cluster. It's cheap, provider-agnostic, and universally supported, though lower-latency-propagation alternatives like global anycast load balancers exist for teams willing to take on that added infrastructure. ## The trade-off on two axes The trade-off against an in-cluster readiness probe is stark on two axes. | Axis | How the two compare | |---|---| | **Propagation latency** | An in-cluster readiness change takes effect in the Service's endpoint list in well under a second; DNS failover takes effect only as fast as caches actually expire across a heterogeneous, uncontrolled population of clients, which in practice is regularly minutes rather than seconds, regardless of the configured TTL. | | **Blast radius** | An in-cluster probe failure affects routing of traffic to one pod; a DNS failure or failover decision affects which of potentially millions of already-cached clients see which region, and because different clients' caches expire at different times, a genuinely split population can end up talking to two different regions concurrently — harmless for stateless, idempotent APIs, but dangerous for anything with regional session affinity, in-flight multi-step transactions, or data that's only eventually consistent across regions. | ## Flapping, and the asymmetric defense Flapping compounds this danger. If a health check oscillates near its failure threshold, because a region is intermittently slow rather than cleanly down, the DNS answer flips back and forth, but because of caching, clients don't cleanly follow the current state; each one sees whatever was cached when its resolver last asked, producing an incoherent client population split unevenly across regions and time. The standard defense is **asymmetric hysteresis**: - require a shorter sustained-failure window to trigger failover, since limiting damage from a real outage matters; - but require a much longer sustained-success window before allowing failback, since a premature failback into a region that's still shaky triggers another round of client splitting almost immediately. The health check driving failover should itself be simple and highly reliable, checked from multiple vantage points, because a false-positive failover is far more expensive and slower to undo than a false-positive pod restart. ## Failure modes that recur at this scale Several failure modes recur specifically at this scale. 1. **Low configured TTLs do not guarantee fast propagation**, because a meaningful fraction of real-world resolvers and client stacks disregard or exceed short TTLs, so observed failover completeness routinely lags the naive TTL-based expectation by minutes. 2. **A 'failback thundering herd'** can occur when many clients' cached answers expire around a similar time, or an automated health check flips the record back the instant its first successful poll comes in, sending a sudden surge of returning traffic into a region that just recovered and hasn't finished warming its own caches or connection pools, echoing the earlier startup-timing problem but at a much larger blast radius. 3. **Regional state splits.** And for any service with regional state, a shopping cart pinned to a region, a live session, a multi-step checkout, clients split across two regions mid-transaction can see inconsistent or lost state, which stateless APIs never have to worry about. ## A documented pattern A documented pattern used with Route 53 failover routing combines a fairly aggressive failover threshold, for instance three consecutive failed checks over ten-second intervals giving roughly thirty seconds to trigger failover, with a deliberately conservative failback approach implemented outside the automatic health check itself — often an on-call engineer manually re-enabling the primary region's record only after confirming full recovery, including cache warm-up and a sustained clean error-rate window, rather than letting the health check flip it back automatically the moment its first successful poll arrives. This asymmetry is precisely what avoids both the flapping and the failback-thundering-herd problems described above.

  • Why can setting a very low DNS TTL, like 5 seconds, fail to actually make failover fast in practice?
    Many real-world resolvers, corporate proxies, certain ISP resolvers, some OS or library stub resolvers, and applications that cache a resolved connection rather than re-resolving per request, either ignore very low TTLs, enforce their own minimum caching duration, or simply don't re-check as often as the TTL suggests. So while a 5-second TTL is a request about how long to cache, actual observed propagation time across a large, heterogeneous client population is often measured in minutes, not seconds.
  • What is a 'failback thundering herd' in the context of DNS-based regional failover, and why does asymmetric hysteresis help?
    It's when a recovered region comes back online and, because many clients' cached DNS answers expire around a similar time or an automated health check flips the record back the instant the region shows healthy, a large burst of previously-diverted traffic returns all at once, potentially overwhelming a region that hasn't finished warming its caches and connection pools. Asymmetric hysteresis, requiring a longer, more conservative sustained-success window before failing back than was required to fail over, gives the recovering region time to stabilize and lets traffic return more gradually or under human supervision.

It's like rerouting mail by changing a forwarding address on file at the post office: everyone who already wrote down your old address keeps using it for a while regardless of what you filed, so some letters go to the old place and some to the new one for days, very different from a receptionist who instantly redirects every call the moment you step out of your office.

saying these in an interview costs you the question

  • Assumes DNS TTL guarantees failover propagates to all clients within that time
  • Treats DNS-based regional failover as equally fast and low-risk as an in-cluster readiness probe
  • Uses symmetric thresholds for failover and failback with no hysteresis
  • Doesn't consider stateful sessions or in-flight transactions when discussing regional split traffic
  • No mention of the failback thundering-herd risk when a region recovers

context