You set up Amazon Route 53 failover records for an active-passive design across two regions. The primary region goes down, but users keep hitting it for several minutes before traffic moves. Walk through everything that contributes to that delay and what you would change.
answer
- detection time plus cache time
- interval times failure threshold
- default is roughly ninety seconds
- TTL keeps the dead answer alive
- lower the TTL before the incident
basics
~20 sTwo delays stack: Route 53 must see enough consecutive failed health-check probes to mark the primary unhealthy, and then every cached copy of the old answer must expire. Shorten both with a faster health check, a lower failure threshold, and a low record TTL.
solid answer
~50 sRoute 53 failover is a detection delay plus a cache delay. Detection is the health check's request interval multiplied by its failure threshold — with the defaults of a 30-second interval and three consecutive failures that is around a minute and a half before the primary is marked unhealthy, plus a little time for the new answer to be served everywhere. Then the cache delay: every resolver and client holding the old answer keeps using it until the record's TTL expires, and clients with open connections may never re-resolve at all. To tighten it I would switch the health check to the 10-second fast interval, consider a threshold of two, and set the failover records' TTL to 60 seconds or lower. I would also be honest that DNS failover has a floor of tens of seconds to minutes — if the recovery objective is tighter than that, the failover mechanism should not be DNS.
go deeper
Know the shape of an active-passive setup: a primary and a secondary record under one name, a health check on the primary, and Route 53 switching answers when the check fails. Recognise that the switch is not instant.
Explain the arithmetic — request interval times failure threshold for detection, then the record's TTL for propagation — and name the defaults. Be ready to say which knobs shorten each part.
Demonstrate production judgment: choose a health-check path that exercises real dependencies, tune interval and threshold against flapping risk, lower TTLs ahead of time, and account for client-side connection reuse. Say out loud what recovery time this design can actually promise.
Own the recovery objective. State the floor DNS failover imposes, decide whether the business requirement justifies moving failover below DNS to anycast or in-region load balancing, and rule on automatic versus human-approved failback given the data layer's replication story.
## How failover records work A failover record set is two records sharing a name and type, distinguished by `SetIdentifier`, with the `Failover` field set to `PRIMARY` on one and `SECONDARY` on the other. Route 53 answers with the primary while it is healthy and with the secondary once the primary is not. The primary must have something telling Route 53 whether it is healthy: either an associated health check (`HealthCheckId`), or — when the record is an alias to an AWS resource such as a load balancer — `EvaluateTargetHealth`, which lets Route 53 inherit that resource's own view of its targets. A point candidates miss: if *both* records are unhealthy, Route 53 does not return an empty answer. It answers as though they were healthy, on the principle that a possibly-broken address beats no address at all. ## Delay one: detection Route 53 runs an endpoint health check from health checkers in multiple AWS regions. Each checker probes on the configured **request interval** — 30 seconds by default, or 10 seconds in fast mode at extra cost — and the endpoint is marked unhealthy after the configured **failure threshold** of consecutive failed probes, default 3 and configurable from 1 to 10. So the default configuration takes roughly 90 seconds of continuous failure before Route 53 changes its mind, and there is a further short interval before that decision is reflected in answers everywhere. An important nuance is that individual checkers disagree — they sit on different networks and see different paths. Route 53 aggregates: the endpoint counts as healthy while a documented minimum share of checkers still report success. That is deliberate. It means one bad network path does not fail your region over, and it also means a *partial* outage — the region is reachable from some networks and not others — may never trip the health check at all. Tuning: fast interval plus a threshold of 2 brings detection to roughly 20 seconds. Do not push it to a threshold of 1 casually; a single lost probe then fails your primary over, and a flapping record is worse than a slow one. ## Delay two: caching Once Route 53 is answering with the secondary, everything holding the old answer still uses it until that copy expires. The record's TTL is the lever you own here: a 300-second TTL means five minutes of continued traffic to a dead region no matter how fast the health check was. Failover records are normally set to 60 seconds or lower for exactly this reason, accepting the extra query volume and cost that a short TTL brings. Two caveats. Alias records to AWS resources do not take a TTL you choose — Route 53 uses a value fixed for the target type — so with an alias to a load balancer the cache delay is largely out of your hands. And the client side is not under your control at all: some runtimes and applications cache resolved addresses beyond the TTL, and an application with an established connection or a warm connection pool will not re-resolve until that connection breaks. This is why a failover that "works" in a test often looks slower in production. ## Making the plan honest Add it up: detection plus TTL plus client behaviour puts realistic DNS failover in the tens of seconds to a few minutes. Design changes that actually help: - **Fast health checks with a sane threshold**, and check an endpoint that reflects the whole stack, not a static file that stays up while the database is gone. A path that touches the real dependencies is what makes the health check mean something. - **Low TTL on the failover records**, set well before you need it — lowering a TTL only takes effect after the old TTL has expired everywhere, so it is not a thing you can do during an incident. - **Make the client cooperate**: sensible connection timeouts and retries so a stuck connection is dropped and re-resolved rather than hanging. - **Fail back carefully.** Automatic failback means the moment the primary passes a couple of probes, traffic returns — possibly to a region that is up but has a cold cache or an unreplicated database. Many teams deliberately make failback a human decision. ## When DNS is the wrong layer If the recovery objective is a few seconds, DNS cannot deliver it, because the caching delay is not yours to control. That is when you move failover below DNS — to anycast addressing that keeps one IP and reroutes the network, or to a load balancer that fails between targets at connection level within a region. Route 53 failover remains the right tool for cross-region, cross-provider, and hybrid failover, where DNS is the only layer that can see both destinations at all.
- Your primary is an alias record pointing at an Application Load Balancer. How does that change the health-check setup?You can set EvaluateTargetHealth on the alias instead of building a separate endpoint health check, and Route 53 then inherits the load balancer's own view of whether it has healthy targets. It saves a health check and reacts to target-level failure. The trade is that you no longer choose the TTL — alias records use a value fixed for the target type.
- Would you set the health check failure threshold to 1 to fail over faster?Rarely. A threshold of 1 turns a single lost probe or a transient network blip into a full regional failover, and flapping between regions is usually more damaging than an extra thirty seconds of downtime — especially with stateful backends. Two failed probes on the fast interval gives roughly twenty-second detection with far more tolerance for noise.
- During an incident you lower the failover record TTL from 300 to 30 seconds. Why does that not help immediately?Because resolvers already cached the record with the 300-second value and will honour it until it expires. The new TTL only applies to answers fetched after the change, so the benefit arrives one old-TTL later. TTL reduction is preparation, not an incident action — lower it in advance of any planned or anticipated cutover.
- Both the primary and the secondary health checks are failing. What does Route 53 answer?It answers as though the records were healthy rather than returning nothing, and for a failover set that means clients still get an address to try. The reasoning is that NODATA guarantees failure while a possibly-working address might not. It also means a total health-check outage — a broken check rather than a broken service — cannot black-hole your domain.
saying these in an interview costs you the question
- Says Route 53 failover happens instantly
- Ignores TTL and cached answers entirely
- Thinks lowering the TTL mid-incident takes effect immediately
- Health-checks a static page that stays up when the app is broken
- Sets the failure threshold to 1 without considering flapping