skip to content

Routing Policies & Health Checks

Route 53's answer-selection logic — weighted, latency-based, failover, geolocation, multivalue — wired to health checks so DNS can steer users away from a sick endpoint. This is the standard mechanism interviewers expect in active-passive and multi-region designs.

part ofAWSoverview, primer and where to startread it →
on this pageshow

questions

6

You set up Amazon Route 53 failover records for an active-passive design across two regions. The primary region goes down, but users keep hitting it for several minutes before traffic moves. Walk through everything that contributes to that delay and what you would change.

level: seniorimportance: must knowfreq 68%

answer

  1. detection time plus cache time
  2. interval times failure threshold
  3. default is roughly ninety seconds
  4. TTL keeps the dead answer alive
  5. lower the TTL before the incident

basics

~20 s

Two delays stack: Route 53 must see enough consecutive failed health-check probes to mark the primary unhealthy, and then every cached copy of the old answer must expire. Shorten both with a faster health check, a lower failure threshold, and a low record TTL.

solid answer

~50 s

Route 53 failover is a detection delay plus a cache delay. Detection is the health check's request interval multiplied by its failure threshold — with the defaults of a 30-second interval and three consecutive failures that is around a minute and a half before the primary is marked unhealthy, plus a little time for the new answer to be served everywhere. Then the cache delay: every resolver and client holding the old answer keeps using it until the record's TTL expires, and clients with open connections may never re-resolve at all. To tighten it I would switch the health check to the 10-second fast interval, consider a threshold of two, and set the failover records' TTL to 60 seconds or lower. I would also be honest that DNS failover has a floor of tens of seconds to minutes — if the recovery objective is tighter than that, the failover mechanism should not be DNS.

go deeper

for a junior

Know the shape of an active-passive setup: a primary and a secondary record under one name, a health check on the primary, and Route 53 switching answers when the check fails. Recognise that the switch is not instant.

for a middle

Explain the arithmetic — request interval times failure threshold for detection, then the record's TTL for propagation — and name the defaults. Be ready to say which knobs shorten each part.

for a senior

Demonstrate production judgment: choose a health-check path that exercises real dependencies, tune interval and threshold against flapping risk, lower TTLs ahead of time, and account for client-side connection reuse. Say out loud what recovery time this design can actually promise.

for a principal

Own the recovery objective. State the floor DNS failover imposes, decide whether the business requirement justifies moving failover below DNS to anycast or in-region load balancing, and rule on automatic versus human-approved failback given the data layer's replication story.

## How failover records work A failover record set is two records sharing a name and type, distinguished by `SetIdentifier`, with the `Failover` field set to `PRIMARY` on one and `SECONDARY` on the other. Route 53 answers with the primary while it is healthy and with the secondary once the primary is not. The primary must have something telling Route 53 whether it is healthy: either an associated health check (`HealthCheckId`), or — when the record is an alias to an AWS resource such as a load balancer — `EvaluateTargetHealth`, which lets Route 53 inherit that resource's own view of its targets. A point candidates miss: if *both* records are unhealthy, Route 53 does not return an empty answer. It answers as though they were healthy, on the principle that a possibly-broken address beats no address at all. ## Delay one: detection Route 53 runs an endpoint health check from health checkers in multiple AWS regions. Each checker probes on the configured **request interval** — 30 seconds by default, or 10 seconds in fast mode at extra cost — and the endpoint is marked unhealthy after the configured **failure threshold** of consecutive failed probes, default 3 and configurable from 1 to 10. So the default configuration takes roughly 90 seconds of continuous failure before Route 53 changes its mind, and there is a further short interval before that decision is reflected in answers everywhere. An important nuance is that individual checkers disagree — they sit on different networks and see different paths. Route 53 aggregates: the endpoint counts as healthy while a documented minimum share of checkers still report success. That is deliberate. It means one bad network path does not fail your region over, and it also means a *partial* outage — the region is reachable from some networks and not others — may never trip the health check at all. Tuning: fast interval plus a threshold of 2 brings detection to roughly 20 seconds. Do not push it to a threshold of 1 casually; a single lost probe then fails your primary over, and a flapping record is worse than a slow one. ## Delay two: caching Once Route 53 is answering with the secondary, everything holding the old answer still uses it until that copy expires. The record's TTL is the lever you own here: a 300-second TTL means five minutes of continued traffic to a dead region no matter how fast the health check was. Failover records are normally set to 60 seconds or lower for exactly this reason, accepting the extra query volume and cost that a short TTL brings. Two caveats. Alias records to AWS resources do not take a TTL you choose — Route 53 uses a value fixed for the target type — so with an alias to a load balancer the cache delay is largely out of your hands. And the client side is not under your control at all: some runtimes and applications cache resolved addresses beyond the TTL, and an application with an established connection or a warm connection pool will not re-resolve until that connection breaks. This is why a failover that "works" in a test often looks slower in production. ## Making the plan honest Add it up: detection plus TTL plus client behaviour puts realistic DNS failover in the tens of seconds to a few minutes. Design changes that actually help: - **Fast health checks with a sane threshold**, and check an endpoint that reflects the whole stack, not a static file that stays up while the database is gone. A path that touches the real dependencies is what makes the health check mean something. - **Low TTL on the failover records**, set well before you need it — lowering a TTL only takes effect after the old TTL has expired everywhere, so it is not a thing you can do during an incident. - **Make the client cooperate**: sensible connection timeouts and retries so a stuck connection is dropped and re-resolved rather than hanging. - **Fail back carefully.** Automatic failback means the moment the primary passes a couple of probes, traffic returns — possibly to a region that is up but has a cold cache or an unreplicated database. Many teams deliberately make failback a human decision. ## When DNS is the wrong layer If the recovery objective is a few seconds, DNS cannot deliver it, because the caching delay is not yours to control. That is when you move failover below DNS — to anycast addressing that keeps one IP and reroutes the network, or to a load balancer that fails between targets at connection level within a region. Route 53 failover remains the right tool for cross-region, cross-provider, and hybrid failover, where DNS is the only layer that can see both destinations at all.

  • Your primary is an alias record pointing at an Application Load Balancer. How does that change the health-check setup?
    You can set EvaluateTargetHealth on the alias instead of building a separate endpoint health check, and Route 53 then inherits the load balancer's own view of whether it has healthy targets. It saves a health check and reacts to target-level failure. The trade is that you no longer choose the TTL — alias records use a value fixed for the target type.
  • Would you set the health check failure threshold to 1 to fail over faster?
    Rarely. A threshold of 1 turns a single lost probe or a transient network blip into a full regional failover, and flapping between regions is usually more damaging than an extra thirty seconds of downtime — especially with stateful backends. Two failed probes on the fast interval gives roughly twenty-second detection with far more tolerance for noise.
  • During an incident you lower the failover record TTL from 300 to 30 seconds. Why does that not help immediately?
    Because resolvers already cached the record with the 300-second value and will honour it until it expires. The new TTL only applies to answers fetched after the change, so the benefit arrives one old-TTL later. TTL reduction is preparation, not an incident action — lower it in advance of any planned or anticipated cutover.
  • Both the primary and the secondary health checks are failing. What does Route 53 answer?
    It answers as though the records were healthy rather than returning nothing, and for a failover set that means clients still get an address to try. The reasoning is that NODATA guarantees failure while a possibly-working address might not. It also means a total health-check outage — a broken check rather than a broken service — cannot black-hole your domain.

saying these in an interview costs you the question

  • Says Route 53 failover happens instantly
  • Ignores TTL and cached answers entirely
  • Thinks lowering the TTL mid-incident takes effect immediately
  • Health-checks a static page that stays up when the app is broken
  • Sets the failure threshold to 1 without considering flapping

context

open as a page

An application runs in three AWS regions behind one name. In Amazon Route 53, when would you choose latency-based routing over geolocation routing, and what does each policy actually decide on?

level: middleimportance: should knowfreq 58%

basics

~20 s

Latency-based routing sends a query to whichever AWS region Route 53 measures as fastest from the requester's network, so it optimises performance. Geolocation routing answers by the requester's mapped location — continent, country or subdivision — so it enforces where traffic must go.

open as a page

Using Amazon Route 53 weighted records, you want roughly 5% of traffic for api.example.com to reach a newly deployed stack. How do weighted records produce that split, and why will the real share of users differ from 5%?

level: middleimportance: should knowfreq 52%

basics

~20 s

Weighted records share a name, and Route 53 picks one per query with probability equal to its weight divided by the total — for example 5 against 95. The real user share drifts because each cached answer serves many clients for the whole TTL.

open as a page

Amazon Route 53's health checkers probe endpoints from the public internet, so they cannot reach an internal load balancer that has only private addresses. How do you still drive Route 53 failover from that endpoint's health?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Use a CloudWatch alarm health check instead of an endpoint health check: Route 53 reads the alarm's state rather than probing anything. Point the alarm at a metric that reflects the private endpoint, such as its unhealthy host count.

open as a page

You are designing multi-region failover on AWS and the target recovery time is a few seconds. Why is Amazon Route 53 failover routing a poor fit for that objective, and what would you do instead?

level: principalimportance: should knowfreq 33%

basics

~20 s

Route 53 failover cannot beat the sum of health-check detection and DNS cache expiry, and the cache half is controlled by resolvers and clients, not by you. For seconds-scale recovery, fail over below DNS — anycast entry points or in-region load balancing.

open as a page

In Amazon Route 53 you can put several IP addresses inside one simple-routing record, or create a set of multivalue answer records for the same name. What does multivalue answer routing give you that simple routing does not?

level: juniorimportance: nice to knowfreq 30%

basics

~20 s

Multivalue answer records each carry their own health check, so Route 53 returns up to eight healthy addresses and omits the failed ones. A simple record always returns all of its values, healthy or not.

open as a page