You are designing multi-region failover on AWS and the target recovery time is a few seconds. Why is Amazon Route 53 failover routing a poor fit for that objective, and what would you do instead?
answer
- detection plus cache, serially
- the cache half is not yours to tune
- changing the answer versus keeping it
- static anycast address, no re-resolution
- routing speed is bounded by data replication
basics
~20 sRoute 53 failover cannot beat the sum of health-check detection and DNS cache expiry, and the cache half is controlled by resolvers and clients, not by you. For seconds-scale recovery, fail over below DNS — anycast entry points or in-region load balancing.
solid answer
~50 sDNS failover has a floor. Route 53 needs several consecutive failed probes to mark the primary unhealthy, and after it switches, every resolver and client holding the old answer keeps using it until the TTL expires — and some clients cache beyond the TTL or never re-resolve while a connection stays open. That puts realistic recovery in the tens of seconds to minutes, and no amount of tuning moves the client-side part because you do not own it. If seconds matter, keep the client's address constant and move the failover below DNS: AWS Global Accelerator gives you static anycast addresses and shifts traffic at the network layer without any name being re-resolved, and within a region a load balancer already fails between targets at connection level. I would also separate the data question from the routing one — routing away in seconds is worthless if the standby's data is minutes behind.
go deeper
Know that Route 53 failover is not instant, because clients keep using cached answers, and that faster failover generally means solving it somewhere other than DNS.
Explain the two serial delays — health-check detection, then TTL-bound caching — and why only the first is genuinely under your control.
Quantify the realistic recovery time a DNS design can promise, name a concrete alternative such as static anycast addressing, and state what the change costs in money and complexity.
Own the whole objective. Bind the routing decision to the data-replication guarantee, decide automatic versus human-gated failover per tier, and be explicit about which failure modes each mechanism actually detects and which it will silently miss.
## Where the floor comes from Route 53 failover is two serial delays. **Detection.** An endpoint health check probes on its request interval and needs a configured number of consecutive failures. Even tuned aggressively — fast interval, threshold of two — that is roughly twenty seconds, and going lower trades reliability for speed because a single lost probe then triggers a regional move. **Propagation.** Route 53 changes what it answers immediately, but the internet is full of copies of the previous answer. Recursive resolvers honour the TTL; some clients cache longer than the TTL; application runtimes and connection pools may hold a resolved address for the process lifetime; and an open, healthy-looking TCP connection to a failing region is never re-resolved at all until it breaks. That second half is the reason DNS cannot deliver seconds. Detection you can tune. Cache behaviour belongs to other people's software, and the only lever you have — the TTL — has to be lowered *before* the incident to have any effect, and buys extra query volume and cost when it is. ## What changing the layer buys The insight to state in an interview is that DNS failover works by **changing the answer**, which means every client must ask again. Faster mechanisms work by **keeping the answer constant** and changing what the address leads to. - **Anycast entry points.** AWS Global Accelerator gives an application two static IP addresses announced from AWS edge locations worldwide. Clients resolve once, possibly forever; when an endpoint group becomes unhealthy, the accelerator steers new connections to another region at the network layer. No re-resolution, no TTL, no client cache in the path. The cost is a service to pay for and operate, and a design constrained to what the accelerator supports. - **In-region load balancing.** Inside a region a load balancer already does connection-level failover between targets, in seconds and without DNS involvement. Plenty of "multi-region" requirements are really availability-zone requirements, and a single regional load balancer across zones satisfies them at a fraction of the complexity. - **Client-side awareness.** Applications with their own retry and endpoint-selection logic — SDKs and mobile clients especially — can move between endpoints faster than any infrastructure mechanism, at the cost of shipping that logic to every client. ## Where Route 53 failover is still exactly right Do not overcorrect. DNS remains the only layer that can see *every* destination, so it is the right instrument when: - the standby is in another cloud, another provider, or an on-premises data centre; - the failover is a whole-stack, deliberate, human-approved event where minutes are acceptable; - the target is a static maintenance page or degraded mode rather than a live standby; - you need geographic or jurisdictional steering as well as failover, which anycast does not express. And Route 53 has a mechanism for the deliberate case that is worth naming: Application Recovery Controller routing controls, an explicit on/off switch per cell that a human or automation flips, with health checks driven by that control rather than by inferred endpoint health. It removes the "did the health check see the right thing?" question from the failover decision, which for a stateful system is often the real risk. ## The part that decides the objective The hardest question in a seconds-scale recovery target is usually not routing at all. If the standby region's database replicates asynchronously, arriving there in five seconds means arriving at data that is behind, and possibly writing into a state that will conflict when the primary returns. A credible design answers three things together: 1. **How traffic moves** — the routing mechanism and the honest recovery time it implies. 2. **What the data guarantees** — synchronous or asynchronous replication, the tolerated data loss, whether writes are single-region or active-active, and how a split brain is prevented. 3. **Who decides** — automatic failover is fast and can fire on a false signal; a human gate is slower but will not shift a stateful system on a network blip. Many mature designs deliberately automate failover for stateless tiers and gate the data tier. A candidate who tunes Route 53 TTLs down to five seconds and calls the objective met has answered the easiest third of the problem. The answer that lands is: DNS is a slow, globally-reachable, coarse control plane; put the fast failover where the client's address does not change; and be explicit that the recovery objective is bounded by the data layer, not by the routing layer.
- Why can't you simply set the TTL to a few seconds and get seconds-scale DNS failover?Because a short TTL only constrains resolvers that honour it. Application runtimes and connection pools cache beyond it, an established connection is never re-resolved, and the query volume — and the bill — rises sharply. It shrinks the average delay without giving a bound, so it improves DNS failover but cannot make it a seconds-scale guarantee.
- What does AWS Global Accelerator change about the failover path compared to Route 53 failover?The client's address stops changing. Global Accelerator publishes static anycast addresses from AWS edge locations, so clients resolve once and the accelerator redirects traffic to a healthy endpoint group at the network layer. No cache expiry sits in the path. The trade is another service to pay for and design around, and less expressive steering than Route 53 policies offer.
- When would you still choose Route 53 failover over a faster mechanism?When the standby is somewhere only DNS can point at — another cloud, another provider, or on-premises — or when the failover is a deliberate, human-approved whole-stack event where minutes are acceptable. Also when the target is a maintenance page rather than a live standby, or when the same records must express geographic steering, which anycast cannot.
- How does the data layer bound the recovery objective?Routing decides where requests land; the data layer decides whether landing there is correct. With asynchronous replication, failing over in seconds means serving data that is behind and risking conflicting writes when the primary returns. The achievable objective is the slower of the two, which is why many designs automate failover for stateless tiers and gate the data tier behind a human decision.
saying these in an interview costs you the question
- Claims a low TTL makes DNS failover near-instant
- Ignores client-side caching and open connections
- Treats routing speed as the whole recovery objective
- Never considers anycast or connection-level failover
- Automates stateful failover without addressing split brain