skip to content

How would you choose a DNS answer-steering design for a service in three regions when its failover speed is bounded by recursive resolver caches you do not control?

level: principalimportance: should knowfreq 18%

answer

  1. decided once per query
  2. set the failover objective first
  3. answers that allow client fallback
  4. anycast for the name servers
  5. fail open, measure from clients

basics

~20 s

Treat DNS steering as locality and minutes-scale evacuation, not seconds-scale failover: anycast the name servers, key answers on ECS where sent, return a local plus an alternate address, fail open, and put faster failover below DNS.

solid answer

~40 s

An authoritative server decides once per query, and afterwards resolvers, hosts and applications hold the answer for as long as they choose, so start from the recovery objective. If it is seconds, DNS alone cannot meet it and something that does not depend on re-resolution must sit in front of the regions. For locality and planned evacuation, combine: anycast authoritative servers so the DNS service is close and survives a site loss; GeoDNS keyed on EDNS Client Subnet where resolvers send it; answers carrying the local region plus a healthy alternate so clients can walk the list (RFC 1123 §2.3); health-checked withdrawal with a probe quorum; and fail-open rather than an empty, negatively cached answer. Then measure where users land and how long drains take, from the client side.

go deeper

for a junior

Know that DNS steering only changes future answers and that cached copies decide how fast clients really move.

for a middle

Describe each steering mechanism, what it keys on, and why anycast name servers do not choose the application address.

for a senior

Design answers that let clients fall back on their own, and a policy for the case where every health check fails.

for a principal

Set the recovery objective first, assign what DNS cannot meet to a lower layer, and argue openly about locality, privacy and automation risk.

## Start from what DNS steering can and cannot control An authoritative server decides once per query it receives. After the answer leaves, recursive resolvers cache it for its TTL, some resolvers send EDNS Client Subnet (RFC 7871) and many do not, hosts may re-sort and hold addresses, and applications may keep connections open for hours. Every design choice below trades against that fact. There is no single right design; the job is to choose which failure modes you accept and to say so. ## The building blocks | Mechanism | What it decides on | Strength | Limit | |---|---|---|---| | Round-robin answers | nothing, order only | clients get alternatives | no locality, load or health | | Weighted answers | a random draw per cache fill | gradual shifts, canaries | exposure follows resolver populations | | GeoDNS | the query's source, usually a resolver | cheap locality | mis-steers users of distant resolvers | | GeoDNS with ECS | the client prefix, where sent | better locality | cache fragmentation, privacy, partial deployment | | Health-checked answers | the checker's state | dead endpoints leave new answers | detection plus cache hold | | Anycast authoritative servers | routing to the nearest instance | fast, resilient DNS service | does not choose the application address | **Anycast name servers** need a separate word. Anycast means one service address made available in several locations, with routing delivering each datagram to one of them (RFC 9499, quoting RFC 4786). Running the zone's authoritative servers this way puts an instance of each name-server address near most resolvers and lets the DNS service survive the loss of a site; resolvers commonly favour name-server addresses with better measured response times, as RFC 1034 §5.3.3 and RFC 1035 §7.2 describe. It does not decide which *application* address is returned: the answer is still keyed on the resolver's address or the ECS prefix. RFC 7766 notes that long-lived TCP to an anycast address can break when routes change, which matters little for short DNS exchanges. ## A decision framework 1. **Write down the failover objective.** Seconds, minutes and "within the hour" lead to different architectures. 2. **Put seconds-scale failover below DNS.** If the objective is a few seconds, DNS cannot meet it alone, because answers already cached outlive any change you make. Use a stable address in front of the regions that does not depend on clients re-resolving, and let DNS handle locality and planned evacuation. 3. **Shape answers for client-side fallback.** Return the local region's address plus a healthy address elsewhere, so a client that walks the list (RFC 1123 §2.3) recovers without waiting for any cache. 4. **Key locality honestly.** Use ECS where resolvers send it, map resolver addresses for the rest, and make every region acceptable, if slower, from anywhere. 5. **Decide the all-unhealthy policy.** Fail open rather than return an empty answer that resolvers negatively cache (RFC 2308). 6. **Treat the TTL as a dial.** Shorter TTLs raise query volume and make you more dependent on your authoritative servers staying reachable; the trade-off itself belongs to the caching topic. Resolvers that implement serve-stale (RFC 8767) may keep answering with expired data when your authoritative servers are unreachable, which helps availability and also prolongs old answers. 7. **Measure from the client side:** where users land, how long a drain takes, how much traffic follows a change, rather than what the policy says. ## Trade-offs to argue out loud - **Locality against privacy and cache cost.** ECS improves steering for users of centralized resolvers, but exposes client prefixes and multiplies cache entries. - **Precision against simplicity.** Weighted shifts are coarse; a precise percentage needs a layer that sees individual users. - **Automation against correlated failure.** An automated checker can evacuate a region in minutes, and a faulty one can evacuate every region at once. Probe quorums, fail-open and a limit on how fast answers may change shrink that blast radius. - **The steering service as a dependency.** If the steering service shares a failure domain with the application, both fail together; keep authoritative service multi-site and independent of the regions it steers. ## An example decision For three regions and a five-minute recovery objective, one defensible design is: anycast authoritative servers; GeoDNS with ECS; each answer carrying the local region plus one healthy alternate; health-checked withdrawal with a quorum of probe locations; fail-open when nothing looks healthy; and a moderate TTL. A tighter objective adds a layer below DNS. A looser one might drop ECS to keep client prefixes private. What makes the answer principal-level is not the list but the reasoning: naming what DNS cannot do, assigning it elsewhere, and choosing the failure you prefer.

  • How do you stop a faulty health checker from evacuating every region at once?
    Require a quorum of probe locations before declaring an endpoint down, fail open when every endpoint looks unhealthy, and cap how many regions the steering logic may withdraw at a time. A checker fault then degrades steering instead of emptying answers, and a real global outage still leaves clients with addresses to try.
  • Does anycasting your authoritative name servers remove the need for GeoDNS?
    No. Anycast decides which name-server instance answers a resolver, which makes lookups faster and more resilient. The answer's content is still chosen by the steering logic from the resolver's address or ECS prefix, so locality of the application address needs GeoDNS or something below DNS.
  • Why return an alternate region's address alongside the local one?
    Because the client can use it without any cache expiring. When the local region fails, a client that walks its address list (RFC 1123 §2.3) connects to the alternate after one connection timeout, while the authoritative withdrawal catches up for new lookups. The cost is that the alternate must always be ready to take that traffic.

saying these in an interview costs you the question

  • A low TTL makes DNS failover as fast as a load balancer
  • Anycast name servers send users to the nearest application region
  • With ECS every user is steered by their own location
  • When all checks fail, return an empty answer
  • DNS steering alone can meet a few-seconds recovery objective