skip to content

Your zone's authoritative DNS servers are unreachable for two hours; what does serve-stale (RFC 8767) let a recursive resolver do, and within what limits?

level: seniorimportance: nice to knowfreq 16%

answer

  1. expired is better than nothing
  2. only when refresh fails
  3. a short TTL on the reply
  4. a cap on staleness

basics

~20 s

Serve-stale (RFC 8767) lets a recursive resolver answer with records whose TTL has expired when it cannot refresh them from the authoritative servers, returning them with a short TTL (30 seconds recommended) and keeping them only for a bounded time.

solid answer

~40 s

RFC 8767 amends the TTL definition: if data cannot be refreshed from its authoritative source when the TTL expires, the resolver MAY use it as though unexpired. When it returns such stale records it MUST give them a TTL above zero, with 30 seconds recommended, so downstream caches re-ask soon. Only an authoritative `NOERROR` or `NXDOMAIN` with the `AA` bit counts as a refresh; timeouts and other rcodes leave the stale data in place, so a deliberate deletion still takes effect. With the zone's own servers down, it only helps names the resolver already had cached, and the RFC's example method keeps expired records for a limited time, 1 to 3 days suggested. It is a resolver's local choice, so an operator cannot count on every resolver doing it.

go deeper

for a junior

Recall that some resolvers can keep answering with an expired record when the servers that own the name cannot be reached.

for a middle

Explain the RFC 8767 rules: stale only after a failed refresh, a TTL above zero on the reply, and which authoritative responses count as a refresh.

for a senior

Reason about an authoritative outage: which names keep working, why a deliberate deletion still wins, and why serve-stale is no substitute for redundant servers.

for a principal

Weigh serve-stale as a cushion you do not control against the cost of redundant authoritative service, and account for it when judging outage impact.

## The problem serve-stale addresses A recursive resolver normally throws away a cached record when its **TTL** (time to live, in seconds) expires, then asks the authoritative servers again. If those servers are down, unreachable or under attack, the lookup fails and clients get `SERVFAIL`, even though the resolver held a perfectly usable answer seconds earlier. For popular names, an outage at the authoritative side becomes an outage for every user as soon as caches expire. **RFC 8767** ("Serving Stale Data to Improve DNS Resiliency", Standards Track) permits a resolver to keep answering with that expired data while it cannot get a fresh answer. ## What the RFC changes RFC 8767 amends the TTL definition in RFC 1035 §3.2.1 and §4.1.3. The new text says: - the TTL is the duration the record **MAY** be cached before the source **MUST** again be consulted; - values **SHOULD** be capped on the order of days to weeks, with a recommended cap of **604800 seconds (7 days)**; - **if the data cannot be authoritatively refreshed when the TTL expires, the record MAY be used as though it is unexpired.** It also changes one detail of RFC 2181 §8: a received TTL with the high-order bit set is now read as a large positive value and capped, instead of being treated as zero. ## The rules for a stale answer 1. **Short TTL on the way out.** A resolver returning stale records **MUST** set their TTL above zero, with **30 seconds RECOMMENDED**. Zero has caused problems in some implementations, and very short values could make TTL-respecting clients retry in a storm; 30 seconds also rate-limits a forwarding resolver behind it. 2. **What counts as a refresh.** A response from an authoritative server with `NOERROR` or `NXDOMAIN` and the **AA** (authoritative answer) bit set **MUST** be treated as refreshing the data. Any other response code **SHOULD** be treated as a failure to refresh, leaving the stale data usable. 3. **Only what is already cached.** When the zone's own authoritative servers are down, a name the resolver never looked up, or whose record it has since evicted, cannot be answered from stale data. (RFC 8767 notes one wider case: stale name-server addresses can still reach a zone whose **parent** is down, even for names not looked up before, provided the zone's own servers are up.) The consequence of rule 2 matters to operators: if you delete a record on purpose and your servers answer NXDOMAIN, serve-stale does **not** keep the old answer alive. The RFC also discusses `REFUSED`: it is ambiguous, but an implementation MAY treat every authority returning it as a signal to stop serving stale data. ## The example method's timers RFC 8767 §5 sketches one way to implement this, with four timers. The values are recommendations or suggestions, not interoperability requirements: | Timer | Purpose | Value the RFC gives | |---|---|---| | Client response timer | How long to try for a fresh answer before falling back to stale data | 1.8 seconds recommended, just under a common 2-second client timeout | | Query resolution timer | Total work spent contacting authorities for one query | commonly around 10 to 30 seconds | | Failure recheck timer | How often to retry failing authorities | no more often than every 30 seconds | | Maximum stale timer | How long an expired record is kept in the cache | suggested 1 to 3 days | The **maximum stale timer** is different from the cap on received TTLs: the 7-day cap clamps how long a record is fresh, while the stale timer bounds how long it lingers **after** expiry. ## What it means for the two-hour outage - Popular names that resolvers had cached keep resolving to their last known addresses; the outage is much smaller than it would otherwise be. - Rarely used names, and names looked up for the first time, fail. - If the outage is exactly the moment you needed to change an address, stale answers point at the **old** one until the servers come back. - Serve-stale is a **local, optional** behaviour. Some resolvers do it and some do not, so a zone operator should treat it as a cushion, not a plan: redundant authoritative servers remain the real defence. ## Why the RFC still calls this responsible The TTL was always a signal for when to refresh, not a command to fail. RFC 8767 keeps that intent: stale data is used only when refreshing has failed, it is handed out with a short TTL so clients come back soon, and as soon as an authoritative answer arrives the stale data is replaced. Its security section notes the edge where this bites: signed data can be returned outside its signature validity period, which is DNSSEC's concern, and only while the authorities are unreachable anyway.

  • Under DNS serve-stale, if an operator deletes a record and the authoritative servers answer NXDOMAIN, does the resolver keep serving the old data?
    No. RFC 8767 says an authoritative `NOERROR` or `NXDOMAIN` with the `AA` bit set must be treated as refreshing the data, so the deletion takes effect. Stale data is only used when the resolver fails to get such an answer, such as timeouts or other response codes.
  • Why does RFC 8767 recommend 30 seconds as the TTL on stale DNS answers rather than zero or the original TTL?
    Zero has caused problems in some implementations, and very short TTLs risk a retry storm from TTL-respecting clients. The original TTL would let downstream caches hold stale data for a long time. Thirty seconds lets them pick up fresh data soon after the authorities recover, while rate-limiting repeat queries.

saying these in an interview costs you the question

  • Serve-stale keeps serving a record even after the operator deletes it.
  • A stale answer goes out with its original TTL restored.
  • Serve-stale lets a resolver answer names it never looked up.
  • Every recursive resolver is required to serve stale data.
  • Serve-stale keeps expired records forever until the servers return.