skip to content

When a DNS authoritative server does health-checked failover for app.example.com, what changes in its answers when an endpoint fails, and what bounds how fast clients follow?

level: seniorimportance: should knowfreq 42%

answer

  1. DNS itself has no health flag
  2. next query, not cached copies
  3. detection time plus cache hold
  4. one address or a list
  5. what if nothing is healthy

basics

~20 s

Health-checked failover changes only what the authoritative server returns to the next queries: the failed endpoint's address is withdrawn or replaced. Clients follow after detection time plus however long resolvers, hosts and applications keep the old answer.

solid answer

~40 s

DNS records have no health state, so failover is steering logic: probes test each endpoint and the authoritative server builds every answer from the endpoints currently passing. In an active-passive design the answer switches from the primary's address to the secondary's; in active-active the failed address drops out of the list. That affects only queries arriving after the checker declares failure. Detection takes several probe intervals, cached answers at resolvers live until their TTL expires, and hosts and applications may hold addresses or connections longer. Answer shape matters: an active-active list lets a client walk to the next address at once (RFC 1123 §2.3). And decide what to return when nothing is healthy: an empty answer is NODATA, which resolvers negatively cache (RFC 2308), so failing open is usually safer.

go deeper

for a junior

Remember that DNS failover only changes future answers, and that cached copies of the old answer keep being used for a while.

for a middle

Explain how the checker's state feeds each answer, and list detection, caches and application behaviour as the parts of the delay.

for a senior

Choose between active-passive and active-active answers, and defend a fail-open policy for the case where nothing looks healthy.

for a principal

Weigh automated DNS failover's reach against the risk that a faulty checker evacuates everything, and decide what must fail over faster than DNS allows.

## What the authoritative server changes DNS itself has no notion of health. An `A` record has no "up" flag, and recursive resolvers do not probe the addresses they cache. **Health-checked failover** is logic in front of, or inside, the authoritative server: probes from one or more vantage points test each endpoint, and the answer-generation step consults the latest health state for every query, building the answer only from endpoints that pass. Some deployments have the checker rewrite the zone instead, for example with DNS UPDATE (RFC 2136); the effect on answers is the same. Two common shapes: | Shape | Answer while healthy | Answer after the primary fails | |---|---|---| | Active-passive | the primary only, e.g. `192.0.2.30` | the secondary only, e.g. `198.51.100.30` | | Active-active | every healthy address | the list without the failed address | The change applies to the **next query the authoritative server receives**, and only to that. An answer already handed out cannot be recalled. ## What bounds how fast clients follow 1. **Detection.** The checker needs several consecutive failed probes before it declares an endpoint down, to avoid flapping on one lost probe. The probe interval and the failure threshold are configuration, not protocol, and together often add up to tens of seconds. 2. **Caches between the server and the client.** Recursive resolvers, forwarders and stub caches keep the old answer until its TTL expires. How to choose that TTL, and why some caches hold answers longer, is the caching topic's subject. 3. **Hosts and applications.** A process may keep a resolved address, or an open connection, well beyond the DNS answer's lifetime. 4. **Client fallback.** Whether a client tries another address after a failed connection is up to the client; RFC 1123 §2.3 says applications SHOULD try multiple addresses until one succeeds. The time from failure until every client has moved is therefore detection, plus the longest cache hold, plus whatever the application does. Only the first term belongs to the steering logic. ## Why the answer's shape matters An **active-passive** answer gives each client exactly one address. When the primary dies, a client holding the cached answer has nothing else to try until the cache expires and a fresh lookup returns the secondary. An **active-active** answer gives clients a list. When one address fails, a client that walks the list connects to another address at once, from the answer it already holds. For such clients failover happens at connection-timeout speed, and the authoritative withdrawal merely stops *new* lookups from including the dead address. The price is that every listed endpoint must be able to take traffic at all times. ## When nothing is healthy If every endpoint fails its checks, or the checker loses its own connectivity and only believes they have, the steering logic must still answer something: - **Return no records.** `NOERROR` with an empty answer section is a **NODATA** response. Resolvers may cache it as a negative answer, and per RFC 2308 its negative TTL is the lesser of the SOA record's own TTL and the SOA `MINIMUM` field. After the endpoints recover, clients behind those resolvers can keep getting no address until that entry expires. - **Fail open.** Return all configured addresses, on the reasoning that a partly working service beats an empty answer and that a fault in the checker is a real possibility. This is a common implementation choice. - **Return `NXDOMAIN`.** Worse than either: it says the name does not exist at all, for every record type, and it is negatively cached too. ## Checker vantage versus user vantage - A checker in one network can see an endpoint as down while users reach it, or the reverse. - Probing from several locations and requiring a quorum reduces false failovers. - A checker's own outage should never be able to empty every answer, which is one more argument for failing open. ## Summary The authoritative server's answers change at the first query after detection; everything after that is bounded by caches and clients the server does not control. Give clients more than one address where you can, decide deliberately what to return when nothing is healthy, and measure failover time from the client's side, not from the checker's log.

  • Why is returning NXDOMAIN a bad choice when every endpoint of a health-checked name looks unhealthy?
    NXDOMAIN (rcode 3) states that the name does not exist at all, for every record type. Resolvers negatively cache it, so the whole name disappears for clients behind them until the negative TTL expires, even after the endpoints recover. Failing open, returning the configured addresses, keeps a partly working service reachable.
  • Can DNS failover move a client's already established TCP connection to the surviving endpoint?
    No. DNS is consulted when a client resolves a name, not while a connection runs. A connection to the failed address breaks or times out, and only when the client reconnects, and resolves again or walks its address list, does it reach the new endpoint.

A shop's phone line that forwards to a branch when the main office closes: callers who dial after the switch reach the branch, but anyone who already wrote the old office's direct number on a note keeps calling it until they look it up again.

saying these in an interview costs you the question

  • Failover is complete once the checker marks the endpoint down
  • Recursive resolvers re-check endpoint health before serving cached records
  • Returning an empty answer when every endpoint is down is harmless
  • DNS failover moves established connections to the new address
  • DNS records carry a health flag that resolvers honour