In an active-active geode deployment - multiple regions, each running a full copy of the backend and application data, any region able to serve any request - how does a client's request typically get routed to a nearby geode, and what happens automatically when one geode's region suffers an outage?
answer
- geo-DNS / anycast / traffic manager
- layered health checks: network, app, synthetic
- no cold standby to warm up
- DNS TTL bounds failover speed
- flapping mitigated with consecutive-failure thresholds
basics
~20 sA smart DNS or traffic-routing service looks at where the request is coming from and sends it to the closest healthy region. If that region's health checks start failing, the router stops sending it traffic and reroutes everyone to the remaining regions instead.
solid answer
~40 sRouting is usually handled by a geo-aware layer sitting in front of all the geodes - options include latency-based or geographic DNS routing, an anycast IP that lets network routing itself pick the nearest point of presence, or a managed global traffic-manager/load-balancer service. That layer continuously runs health checks (often at multiple levels: network reachability, application-level liveness, sometimes a synthetic transaction) against each geode. When a geode's health checks fail, the router marks it unhealthy and stops directing new traffic to it, redistributing that traffic among the remaining healthy geodes based on proximity or weighted priority - all without a human triggering a failover, because every remaining geode is already running and already has the data to serve those requests.
go deeper
Should know, at a high level, that some kind of smart routing sends users to a nearby region and that unhealthy regions get automatically skipped. Doesn't need to name specific mechanisms like anycast or BGP.
Should be able to name at least one concrete routing mechanism (geo-DNS, traffic manager, anycast) and describe that health checks continuously monitor each geode so unhealthy ones are removed from rotation without manual intervention.
Should understand the trade-offs behind TTL tuning and health-check sensitivity, including flapping and hysteresis, and be able to reason about the gap between 'is up' and 'is healthy for real traffic.'
Should be able to design a layered health-check strategy (network/app/synthetic), justify TTL and threshold choices against business latency/availability requirements, and anticipate second-order failure modes like health-check dependency loops causing cascading regional evacuations.
## The layer in front of the fleet Getting a request to the right geode, and getting it away from a broken one, is the job of a **routing/health-check layer** that sits logically in front of the fleet of regional deployments. There are a few concrete mechanisms teams use for the routing half. - **Latency-based or geographic DNS** steers a client to a region based on either the client's resolved location or measured network latency from various vantage points, resolving the same hostname to different regional IP addresses depending on who's asking. - **Anycast** takes a different approach: the same IP address is announced from multiple regions, and ordinary internet routing (BGP) delivers each packet to whichever announcing location is topologically nearest, so 'nearest geode' falls out of the network layer itself rather than DNS. - **Many teams instead use a managed global traffic-manager or global load-balancer product** that offers a hybrid of geo-proximity, latency, and weighted routing policies, plus built-in health checking, so they don't have to hand-roll DNS logic. Whichever mechanism is chosen, the routing decision usually isn't purely static - it's continuously informed by health signals, which is the second half of the picture. ## Health checks come in layers Health checking in this context is layered. - **A basic check** might just confirm the region's load balancer or network path is reachable. - **A more meaningful check** calls an application-level health endpoint that verifies the service process is up and can reach its own dependencies (database connection, cache, downstream services). - **The most thorough setups run synthetic transactions** - a scripted request that exercises a real code path - on a schedule, so that a geode which is technically 'up' but returning errors or badly elevated latency for real traffic still gets flagged. These checks run continuously and independently of any actual outage, which is what makes automatic failover possible: the system already knows a geode is unhealthy within the health-check interval (commonly tens of seconds), not after users start complaining. ## What an evacuation actually looks like When a geode's checks start failing, the routing layer marks it out of rotation and stops sending new requests there. Existing in-flight requests to that geode still fail or time out - a real user-facing blip, not a magic instant fix - but new requests get routed to the remaining healthy geodes, chosen the same way, by proximity, latency, or configured weight. Because every geode was already running a full copy of the application and already held a replicated copy of the data, this reroute is a matter of **DNS TTL expiry** or **traffic-manager reconvergence**, seconds to low minutes depending on the mechanism and configured TTLs, not a cold start of new infrastructure. That's the central operational payoff of active-active over active-passive: there's no 'promote the standby, warm its caches, redirect traffic' runbook to execute under pressure, because the standby doesn't exist - every region was already a full peer. ## The edges of the automation The trade-off is that this automation has real edges. - **DNS-based routing is bounded by how aggressively clients and resolvers honor TTLs** - a very low TTL improves failover speed but increases DNS query volume and cost, while a high TTL means some clients keep hammering a dead region for longer than the health check alone would suggest. - **Anycast reconvergence depends on how quickly BGP routes propagate**, which is mostly out of the application team's control. - **Health checks themselves can be wrong in both directions**: too shallow a check (just 'is the load balancer up') can leave a genuinely broken geode in rotation because the check doesn't exercise the failing code path, while too aggressive a check (flagging on transient latency blips) can cause 'flapping,' where a geode gets pulled in and out of rotation repeatedly, adding instability instead of removing it. Teams typically tune this with **hysteresis** - requiring several consecutive failed checks before pulling a geode out, and several consecutive successes before adding it back - to avoid flapping. ## Failure modes Failure modes commonly seen in production: - a **'silent brownout,'** where a geode is slow or erroring for a subset of endpoints but the shallow health check still passes, so traffic keeps flowing there and users in that region see a degraded experience with no automatic mitigation; - a **health-check dependency loop**, where the health endpoint itself calls a downstream service that's degraded, causing the check to fail (and traffic to reroute) even though the core service is actually fine, amplifying a small dependency problem into a full regional evacuation; - and **DNS-caching issues** at the client or corporate-proxy level that keep sending some fraction of users to an already-evacuated region well past the intended TTL. ## Where it shows up A concrete real-world pattern: large content and API platforms commonly pair a global traffic-manager or anycast-fronted CDN/API gateway with regional backend geodes, configuring latency-based routing plus multi-level health checks (network, app, and synthetic) with a short but not razor-thin TTL and a few-consecutive-failure threshold before evacuating a region - trading a little failover speed for protection against flapping on transient blips.
- What's the practical downside of setting DNS TTLs very low to speed up failover?A very low TTL means resolvers re-query far more often, which increases DNS query volume, load on the DNS infrastructure, and sometimes cost, and it can also increase perceived latency slightly since more requests pay a DNS lookup cost instead of using a cached answer. It's a genuine trade-off between failover speed and DNS overhead, so teams tune it rather than setting it as low as technically possible.
- Why do teams require several consecutive failed health checks before pulling a geode out of rotation, instead of reacting to the very first failure?A single failed check is often just noise - a transient network blip or a momentary GC pause - and reacting to it immediately can cause the geode to flap in and out of rotation, which is more disruptive than a brief real outage would be. Requiring a few consecutive failures, and a few consecutive successes to rejoin, adds hysteresis that filters out noise while still reacting to genuine sustained problems within a bounded time.
- How is anycast-based routing different from DNS-based geo-routing?DNS-based routing makes the routing decision at name-resolution time by returning different IP addresses to different clients based on location or measured latency, so the client ends up with one specific regional address. Anycast instead announces the same IP address from every region and lets normal internet routing deliver each packet to whichever announcing location is network-topologically closest, so the 'nearest region' decision is made continuously by the network layer rather than once at DNS lookup time.
Like a GPS app rerouting you away from a road it detects is jammed - it's constantly checking conditions, and the moment one route looks bad it quietly sends new traffic down the other roads that were already open, without you having to build a new road first.
saying these in an interview costs you the question
- Believes failover requires someone to manually promote a standby region
- Assumes DNS changes propagate instantly with no TTL delay
- Only mentions checking that the load balancer responds, with no application-level or synthetic health check
- Doesn't recognize that overly sensitive health checks cause flapping
- Thinks anycast and geo-DNS are the same mechanism