skip to content

Caching and TTL

Every hop caches an answer for the TTL the zone sets, negative answers included. Lowering TTLs before a migration and 'why is my change not live yet' are the angles interviewers reach for.

on this pageshow

questions

5

In DNS, what does a record's TTL control, and why can clients keep getting the old address after you change the record?

level: juniorimportance: must knowfreq 68%

answer

  1. set by the zone's owner
  2. seconds, not hops
  3. every cache counts down
  4. old copies live until expiry

basics

~20 s

A DNS TTL is the number of seconds a cache may keep a record before consulting the authoritative source again. Resolvers that cached the old record can keep serving it until their copy's remaining TTL runs out.

solid answer

~50 s

The TTL is a 32-bit field on every DNS resource record, set by the zone's administrator, giving the number of seconds a resolver may cache that record before it must ask the source again (RFC 1035 §3.2.1; RFC 2181 §8 limits the value to 0 through 2^31-1). Changing the record on the authoritative server pushes nothing out: every recursive resolver that already fetched the old answer may keep serving it until its copy expires, so the change appears gradually over up to one old TTL. A cache hands out the *remaining* TTL, counted down since it fetched the record, so a chain of conforming caches does not stretch that window. A TTL of `0` means use the record for the current transaction only. It is a cache lifetime in seconds and has nothing to do with the IP header's TTL, which counts router hops.

go deeper

for a junior

Recall that a TTL is the number of seconds a cache may keep a record, set by the zone owner, and that already cached copies survive until they expire.

for a middle

Explain the countdown: each cache hands on the remaining TTL, so the spread of a change is bounded by the old TTL rather than multiplied by the number of caches.

for a senior

Show that the TTL is a ceiling rather than a floor (RFC 2181 §8), that holders outside the resolver can ignore it, and that change plans start from the old value.

for a principal

Treat the TTL as a lever between change agility and query load or resilience, and set it per record according to how often that record must move.

## What the TTL field is Every DNS **resource record** (an A record holding an IPv4 address, an AAAA record holding an IPv6 address, an MX record naming a mail server, and so on) carries a **TTL** (time to live). RFC 1035 §3.2.1 defines it as the time interval, in seconds, that the record **may be cached before the source of the information should again be consulted**. - The value is chosen by whoever administers the **zone** the record lives in, not by the resolver that caches it. - RFC 1035 described the field inconsistently (signed in one place, unsigned in another); **RFC 2181 §8** fixed it as an unsigned number from `0` to `2147483647` (2^31-1). - RFC 2181 §8 also says the TTL is a **maximum time to live, not a mandatory one**: a cache may throw a record away earlier, and may cap very large values. - **RFC 8767** later recommended that resolvers cap TTLs at `604800` seconds (7 days). - All records of one type at one name (an **RRset**) must share the same TTL (RFC 2181 §5.2). - A TTL of `0` means the record can be used only for the transaction in progress and should not be cached. ## Who counts it down A DNS answer usually passes through more than one cache before an application uses it. The authoritative server holds the zone itself; a **recursive resolver** fetches and caches answers for many clients; a forwarding resolver or a cache on the client side may sit in front of it. | Holder | What it keeps | TTL it hands on | |---|---|---| | Authoritative server | The zone data | The full TTL written in the zone | | Recursive resolver | A copy fetched at some moment | The **remaining** TTL: original minus time already cached | | Downstream cache | A copy of the resolver's copy | Whatever is left of the remaining TTL | Because each conforming cache passes on only what is left, the copy's expiry time stays fixed as it moves downstream: a chain of caches does **not** add their delays together. RFC 1034 shows the telltale sign of a cached answer: the **AA** (authoritative answer) bit is not set and the TTL is lower than the zone's value, the difference being the time the data has aged in the cache. ## Why a change is not live everywhere at once Editing the record on the authoritative server is not a push. Nothing tells the caches that already hold the old answer. Consider a record with TTL `3600` changed at 10:00: 1. Resolver A fetched the old record at 09:00. Its copy expires at 10:00, so it fetches the new answer on its next lookup. 2. Resolver B fetched the old record at 09:50. Its copy is valid until 10:50, and it serves the old address for the rest of that hour. 3. Resolver C had never looked the name up. Its first lookup after 10:00 gets the new answer immediately. The result is a gradual rollover lasting up to **one old TTL**. "DNS propagation" is really this: independent caches expiring at different moments. Flushing the cache on your own machine only fixes your own view; other resolvers' copies are outside your control. Two consequences follow: - The TTL that governs how fast a change spreads is the **old** one, the value that caches received when they last fetched the record. A lower TTL published now only reaches a cache when its current copy expires. - Some holders go beyond what the resolver does: an application or long-running process may keep its own copy of an answer, or reuse an open connection, regardless of the TTL. The DNS specification bounds the resolver caches, not those. ## Two TTLs that are not the same The word **TTL** appears in two unrelated places: - The **DNS TTL** is a cache lifetime in **seconds**, carried inside each resource record. - The **IP TTL** (the hop limit in IPv6) is a field in the IP header, decremented by each router, that stops a packet looping forever. It counts **hops** and has nothing to do with how long a DNS answer is cached. Confusing the two is a common interview slip: a DNS reply crossing twelve routers does not change the record's TTL at all. ## What to take away A TTL is a promise the zone owner makes to every cache: this answer stays usable for this many seconds. It trades freshness against load: a long TTL means fewer queries and faster lookups but slower changes, a short TTL means faster changes but more queries to the authoritative servers. When someone asks why a DNS change "has not propagated", the answer is almost always that caches are still inside the old TTL.

  • If a DNS record has a TTL of 0, is every lookup guaranteed to reach the authoritative server?
    No. RFC 1035 says a zero TTL means the record should not be cached and is usable only for the transaction in progress, and conforming resolvers follow that. But it is a 'should', and caches outside the resolver, such as an application holding its own copy, are not bound by it. It also multiplies query load and adds a full lookup to every request.
  • How can you tell from a DNS response whether it came from a cache rather than the authoritative server?
    A cached answer does not have the `AA` (authoritative answer) bit set, and its TTL is lower than the value in the zone because the resolver hands out the remaining lifetime of its copy. RFC 1034 uses exactly that difference to show an answer that was served from a cache.

A DNS TTL works like a printed timetable stamped 'valid until' a date: people who picked up a copy keep using it until that date even if the master timetable changes, and a photocopy made later carries the same expiry date, not a fresh one.

saying these in an interview costs you the question

  • The DNS TTL is the number of routers a DNS packet may cross.
  • Changing the record on the authoritative server pushes the update to resolvers.
  • Each cache in a chain restarts the full TTL, so delays add up per hop.
  • Flushing my own machine's DNS cache makes the change live for everyone.
  • A lower TTL published now shortens copies that are already cached.
open as a page

You must move www.example.com, whose DNS A record has a TTL of 86400, to a new server with minimal stale traffic; how do you schedule the TTL changes?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Lower the TTL (say to 300) at least one old TTL, 86400 seconds, before the switch, so every long-lived copy expires first. Then change the address, keep the old server answering briefly, and raise the TTL back once stable.

open as a page

A DNS name was queried before its record existed, and resolvers keep returning NXDOMAIN after you create it; what sets how long?

level: middleimportance: should knowfreq 38%

basics

~20 s

Negative caching (RFC 2308): the NXDOMAIN reply carries the zone's SOA record, whose TTL is the smaller of the SOA's own TTL and its MINIMUM field. Resolvers keep answering NXDOMAIN until that negative TTL expires.

open as a page

How do you choose DNS TTLs for a service that relies on DNS-based failover, and what do very short TTLs cost you?

level: principalimportance: should knowfreq 22%

basics

~20 s

Worst-case DNS failover time is detection plus publishing plus one full TTL, so failover names need short TTLs. Short TTLs cost query load, lookup latency and dependence on authoritative uptime, so set them per record, not zone-wide.

open as a page

Your zone's authoritative DNS servers are unreachable for two hours; what does serve-stale (RFC 8767) let a recursive resolver do, and within what limits?

level: seniorimportance: nice to knowfreq 16%

basics

~20 s

Serve-stale (RFC 8767) lets a recursive resolver answer with records whose TTL has expired when it cannot refresh them from the authoritative servers, returning them with a short TTL (30 seconds recommended) and keeping them only for a bounded time.

open as a page