skip to content

You must move www.example.com, whose DNS A record has a TTL of 86400, to a new server with minimal stale traffic; how do you schedule the TTL changes?

level: seniorimportance: must knowfreq 52%

answer

  1. caches keep the TTL they got
  2. lead time equals the old TTL
  3. switch while short, then raise
  4. keep the old server answering

basics

~20 s

Lower the TTL (say to 300) at least one old TTL, 86400 seconds, before the switch, so every long-lived copy expires first. Then change the address, keep the old server answering briefly, and raise the TTL back once stable.

solid answer

~40 s

Caches hold the TTL they received, so a lowered TTL only reaches a cache when its current copy expires. I would publish TTL `300` at least 86400 seconds before the cutover, plus a margin, because a resolver that fetched the record just before the change keeps the 86400-second copy for a full day. At cutover I change the A record; conforming caches then move to the new address within 300 seconds. I keep the old server answering for longer than that, because some holders ignore TTLs. While the TTL is short, rollback is equally fast, so I keep it short until I am sure, then raise it back to its old value, which RFC 1034 §3.6 describes as the expected practice.

code

dns · 8 lines
dns
; T-24h and earlier: lowered ahead of the move
www.example.com.  300    IN  A  192.0.2.10

; T: the cutover, still short so rollback is fast
www.example.com.  300    IN  A  198.51.100.20

; after it has proved stable: raised back
www.example.com.  86400  IN  A  198.51.100.20

go deeper

for a junior

Recall that caches keep the TTL they fetched, so a short TTL has to be published well before the address changes.

for a middle

Compute the lead time from the old TTL, not the new one, and explain why the new TTL then bounds the stale window after the switch.

for a senior

Plan the whole window: lead time, cutover, keeping the old server answering for holders that ignore TTLs, the rollback path and the moment to raise the TTL.

for a principal

Set per-record TTL policies so routine moves do not need a day of preparation, weighing the query load and resilience that short TTLs cost.

## The rule that drives the plan A DNS **TTL** is the number of seconds a cache may hold a record before asking the authoritative source again. The crucial detail for a migration is that a cache keeps the TTL it received **when it fetched the record**. Publishing a new, lower TTL on the authoritative server does nothing to copies already sitting in resolvers; each of those copies runs out on its own schedule. RFC 1034 §3.6 states the practice directly: if a change can be anticipated, the TTL can be reduced prior to the change to minimise inconsistency, and then increased back to its former value after the change. ## Working the numbers The record is `www.example.com` with TTL `86400` (one day). The target TTL during the move is `300` (five minutes); the choice of 300 is an operational judgement, not an RFC value. 1. **T-24h minus a margin: lower the TTL to 300.** A resolver that fetched the record one second before this edit holds an 86400-second copy. It will not ask again for a full day. 2. **Wait at least 86400 seconds.** After one old TTL, every conforming cache that held a long copy has expired it and re-fetched a copy with TTL 300. 3. **T: change the A record** from the old address to the new one. 4. **T + 300 s:** every conforming cache has expired its last copy of the old address and fetched the new one. 5. **Hours to days later: raise the TTL back** once traffic on the new server looks right and you no longer expect to roll back. The **minimum lead time equals the old TTL**, not the new one. Lowering the TTL an hour before a switch, when the old TTL was a day, leaves most of the stale window in place. | Time | Action | Worst stale window for a conforming cache | |---|---|---| | Before T-24h | TTL 86400 | up to 86400 s | | T-24h (minus margin) | TTL lowered to 300 | existing long copies still valid | | T | Address changed | up to 300 s | | After stabilising | TTL raised back | irrelevant: no change in flight | ## Why keep the old server answering The TTL bounds what conforming resolver caches do. It does not bound: - an application or long-running process that resolved the name once and keeps using the address, or keeps an open connection; - a resolver configured with its own **minimum cache TTL**, an implementation choice that holds records longer than the zone allows and exceeds the maximum that RFC 2181 §8 defines; - clients that were already mid-session when the switch happened. So the old server should keep answering, or forward to the new one, until its traffic has fallen to nothing that matters. Watching that traffic decay is also the evidence that the cutover worked. ## Rollback and when to raise the TTL While the TTL is 300, a rollback spreads in five minutes too. That is the reason to keep it short until the new server has proved itself, and only then raise it. Raising it is not instant either: caches pick up the longer value as they re-fetch, which is harmless. A TTL of `0` for the migration window is rarely worth it. RFC 1035 says a zero TTL should not be cached, so every client lookup has to reach the authoritative servers: query volume and lookup latency climb, and it buys little over a few minutes of TTL. ## Things that change the plan - **CNAME chains.** Each RRset in a chain is cached separately with its own TTL. Lower the TTL of the record whose data will actually change: the CNAME if you repoint it, or the target's A record if the target moves, which may belong to someone else's zone. - **New names.** If the move introduces a name that did not exist before, anyone who queried it early may hold a cached negative answer whose lifetime comes from the zone's SOA record, not from the record you create. - **Moving the name servers themselves** involves the parent zone's delegation and its TTLs, a separate problem from the A record. ## What the interviewer is listening for The weak answer is "lower the TTL at the moment of the switch". The strong answer names the old TTL as the lead time, the new TTL as the stale window, the old server as the safety net for holders that ignore TTLs, and the raise-back as a deliberate final step.

  • Why not set the DNS record's TTL to 0 for the migration window instead of 300?
    RFC 1035 says a zero TTL means the record should not be cached, so every client lookup has to reach the authoritative servers. That multiplies query load and adds a full lookup's latency to every request, while a TTL of a few minutes already bounds the stale window closely enough. Some holders ignore TTLs anyway, so zero does not buy certainty.
  • www.example.com is a CNAME to a name in another zone. Which DNS TTL do you lower before the move?
    Each RRset in the chain is cached on its own with its own TTL, so lower the one whose data will change. If you repoint the CNAME to a new target, lower the CNAME's TTL. If the target keeps its name but its A record changes, it is the target's TTL that matters, and that zone's owner controls it.
  • When is it safe to raise the DNS TTL back after the cutover?
    Once traffic has moved, the new server has run long enough that you no longer expect to roll back, and residual traffic at the old server is negligible. Raising it early gives up the five-minute rollback. The raise itself spreads gradually as caches re-fetch, which is harmless.

saying these in an interview costs you the question

  • Lower the TTL at the moment of the switch and the change is live in minutes.
  • Lowering the TTL makes resolvers drop their existing cached copies at once.
  • After the short TTL has passed, nothing can still reach the old address.
  • A TTL of 0 for the migration is free and the safest choice.
  • Keep the TTL short permanently afterwards; short TTLs cost nothing.