How do you choose DNS TTLs for a service that relies on DNS-based failover, and what do very short TTLs cost you?
answer
- failover time includes the TTL
- load scales with one over TTL
- short TTLs lean on authoritative uptime
- ceiling, not a guarantee
basics
~20 sWorst-case DNS failover time is detection plus publishing plus one full TTL, so failover names need short TTLs. Short TTLs cost query load, lookup latency and dependence on authoritative uptime, so set them per record, not zone-wide.
solid answer
~50 sWith DNS-based failover, the time until conforming caches move is roughly detection time plus the time to publish the change plus one full TTL, because a cache that refreshed just before the switch keeps the old answer for its whole TTL. So the records that move need short TTLs, often tens of seconds to a few minutes; that is an operational choice, not an RFC value. The costs: authoritative query volume grows roughly as one over the TTL, more lookups miss the cache and pay a full resolution, and the name disappears quickly if the authoritative servers go down, unless resolvers serve stale data. I keep short TTLs on the few steerable front-door records and long ones on stable records like `NS` and `MX`, and I treat the TTL as a ceiling for conforming caches, not a promise about every client.
go deeper
Recall that a DNS change only reaches users as caches expire, so a record meant to fail over quickly needs a short TTL.
Explain the failover bound as detection plus publication plus one TTL, and why a chain of conforming caches does not lengthen it.
Quantify the cost of short TTLs in authoritative query load and latency, and account for holders that ignore TTLs when promising failover times.
Set TTLs per record class against the resilience trade-off with authoritative availability, and decide when DNS is the wrong layer for failover altogether.
## What bounds DNS failover time In DNS-based failover, a health check notices a failed endpoint and the authoritative answer for a name is changed to point elsewhere. How quickly clients move is bounded by caching: 1. **Detection**: the health check must decide the endpoint is down. 2. **Publication**: the authoritative servers must start giving the new answer. 3. **Cache expiry**: a resolver that refreshed its copy just before the change keeps the old answer for its **full TTL** (time to live, in seconds). So the worst case for conforming caches is roughly **detection + publication + TTL**. With a 60-second detection and a TTL of 300, clients behind a resolver that has just refreshed keep reaching the dead endpoint for about six minutes. A chain of conforming caches does not make this longer, because each passes on only the remaining TTL. Two limits sit outside that formula: - RFC 2181 §8 makes the TTL a **maximum**, so caches may refresh sooner, but some holders keep answers longer: applications with their own copy, open connections that DNS never touches, or resolvers configured with a minimum cache TTL. - DNS failover moves **new** lookups. It does nothing for clients already connected to the failed endpoint. ## What short TTLs cost | Effect | Long TTL (hours to days) | Short TTL (seconds to minutes) | |---|---|---| | Failover and rollback speed | slow: up to the whole TTL | fast | | Query load on authoritative servers | low | high: roughly proportional to 1/TTL | | Lookup latency for users | mostly cache hits | more cache misses and full resolutions | | Survival of an authoritative outage | caches carry the name for hours | names vanish within one TTL, unless resolvers serve stale data | The load arithmetic is rough but useful. A busy resolver refreshes an actively used name at most about once per TTL. If 10,000 resolvers are actively using a name, a TTL of 3600 means on the order of 2.8 refresh queries per second to the authoritative servers; a TTL of 60 means about 167 per second, sixty times more, for the same users. The resilience cost is easy to miss. With long TTLs, an outage or denial-of-service attack on the authoritative servers hurts only new lookups for a while. With 30-second TTLs, every cache loses the name within half a minute. **RFC 8767** serve-stale softens this, since a resolver may keep returning expired data when refresh fails, but it is optional and a resolver's local choice, so you cannot plan on it. ## A decision frame There is no single right number; the question is which records need to move and how fast. - **Classify records by how they change.** `NS` and `MX` records and the addresses of the name servers themselves change rarely; hours to a day fits them. RFC 1034 §3.6 recommends TTLs on the order of days for typical hosts, lowered only around anticipated changes. - **Short TTLs only where failover is the design.** The front-door `A`/`AAAA` records that a health check rewrites are the candidates, typically tens of seconds to a few minutes. Those numbers are operational practice, not RFC requirements. - **Size the authoritative service for the short TTLs you choose**, with redundant, well-distributed servers, because you have made users depend on them every minute. - **Remember the negative TTL.** The zone's SOA record sets how long a "does not exist" answer is cached (RFC 2308). A failover design that briefly removes a name, or creates a new one, is at the mercy of that value. - **Measure, do not assume.** After a failover, traffic keeps arriving at the old endpoint from holders that ignore TTLs; watching that tail tells you what failover time you really get. ## When DNS is the wrong lever If the requirement is failover in seconds for every client, including connected ones, DNS caching is working against you. Designs that keep one stable address and move traffic behind it, such as anycast or a load-balancing layer, avoid the cache entirely; those belong to other parts of the architecture. DNS failover suits cases where minutes of partial impact are acceptable. ## What a strong answer shows It states the failover bound with the TTL in it, quantifies the query cost, names the resilience trade-off with authoritative availability, sets TTLs per record class rather than zone-wide, and is honest that the TTL bounds conforming caches rather than every client.
- Your DNS health check fails over in 30 seconds and the record's TTL is 60. What failover time do you promise, and why not 90 seconds flat?About 90 seconds plus publication time is the bound for conforming caches that refreshed just before the switch. I would not promise it for all clients: some applications and resolvers hold answers beyond the TTL, and existing connections to the failed endpoint are unaffected by DNS. I would measure the residual traffic tail and quote that.
- Why can very short DNS TTLs make an attack on your authoritative servers more damaging?With short TTLs every resolver's copy expires within seconds or minutes, so once the authoritative servers stop answering the names disappear for almost everyone. Long TTLs let caches carry the name through the outage. Resolvers that implement RFC 8767 serve-stale soften this, but it is optional, so it cannot be relied on.
saying these in an interview costs you the question
- A short TTL guarantees every client fails over within that many seconds.
- Short TTLs are free, so set every record in the zone to 30 seconds.
- DNS failover moves existing connections off the failed endpoint.
- Several caches in a row multiply the failover delay by the number of caches.
- TTL choice has no bearing on how an authoritative outage affects users.