skip to content

A DNS record was repointed an hour ago. A fresh lookup on a Linux host returns the new address, but a long-running service on that same host keeps connecting to the old one. Where can the stale address still be held?

level: seniorimportance: nice to knowfreq 33%

answer

  1. the lookup itself is not the problem
  2. several layers may hold a copy
  3. the C library keeps nothing
  4. a stub daemon, or the process itself
  5. an open connection never re-resolves

basics

~20 s

Not in the C library, which caches nothing between calls. Look at a local caching resolver on the host, an NSS caching daemon, the process's own in-memory address cache, and an already-established connection that was opened to the old address and never closed.

solid answer

~50 s

Rule out the C library first: the glibc resolver performs a query on every call and keeps no cache, so a fresh lookup returning the new address tells you the resolution path is fine. The stale value therefore lives in one of three other places. A **local caching resolver** such as systemd-resolved holds answers on the host — `resolvectl flush-caches` clears it and `resolvectl statistics` shows whether it is being hit. An **NSS caching daemon**, where one is still deployed, can hold host entries independently. Most often it is the **process itself**: many runtimes and clients cache resolved addresses in memory, sometimes for the process's lifetime. And a fourth case is not a cache at all — a **connection already established** to the old address, or one pooled and kept alive, will keep using it until it is closed, no matter what DNS now says.

code

bash · 2 lines
bash
resolvectl statistics
resolvectl flush-caches

go deeper

for a junior

Know that a name is resolved when a connection is opened, not continuously, and that a running program may keep using an address it looked up earlier.

for a middle

Explain that the glibc resolver itself caches nothing, so a stale answer must come from a caching daemon on the host or from the application's own memory, and know how to inspect and clear the local resolver cache.

for a senior

Diagnose by elimination: prove the host's lookup is correct, then distinguish a local cache from an in-process cache from an established connection, and pick the action that actually applies rather than flushing everything.

for a principal

Own the failover design implication: a name change is best-effort and its worst case is the slowest cache plus the longest-lived connection. Decide where cutovers should be steered and how connection lifetimes and client caches are bounded fleet-wide.

## Start by ruling out the resolver The useful first fact is a negative one: **the glibc resolver does not cache**. Every `getaddrinfo()` call performs the lookup again. So if a fresh lookup on the host returns the new address, the host's resolution path is already correct and there is no point flushing anything at that layer. The stale value is being held somewhere further along, and the job is to work out where. ## Layer by layer **A local caching resolver.** If `/etc/resolv.conf` points at a loopback address, a daemon on the host is caching answers. systemd-resolved keeps both positive and negative entries; `resolvectl statistics` reports cache hits and `resolvectl flush-caches` empties it. This layer is invisible to anyone who assumes Linux "has no DNS cache", which is only true of the C library. ```bash resolvectl statistics resolvectl flush-caches ``` **An NSS caching daemon.** Historically `nscd` cached hosts entries behind the switch interface, and directory-integration daemons can cache too. It is far less common on current systems — recent glibc releases have dropped nscd — but on an older or specially configured host it is a real place for an answer to hide, and it caches at the NSS layer, so even `getent` can return a stale value. **The process's own memory.** This is the usual culprit for the symptom as described. Many language runtimes and HTTP client libraries resolve a name once and keep the result, some for a configurable interval and some for the lifetime of the process. Nothing on the host can flush that; you either configure the client's own cache lifetime or restart the process. **An open connection.** The fourth case is not caching at all, and it is the one candidates most often miss. DNS is consulted only when a *new* connection is opened. A long-lived TCP connection, or an idle connection kept in a client pool for reuse, was established to the old address and stays there until something closes it. The service is not resolving anything wrongly — it is not resolving at all. ## How to tell them apart Work outside in, and let each test eliminate a layer: 1. Do a fresh lookup on the host. New address? The resolution path is fine — stop debugging DNS. 2. Query an upstream server directly and compare with a lookup through the local stub. A difference means a local resolver cache. 3. Compare through the NSS path as well; if that disagrees with a direct DNS query, look at the switch configuration and any caching daemon behind it — and check `/etc/hosts` for a pin someone added during an earlier incident and forgot. 4. Look at the peer addresses the process is actually connected to. If it is holding established connections to the old address, no cache is involved; the connections must be closed. 5. If everything on the host is correct and the process still misbehaves, the cache is inside the process. Restarting it is the decisive test. ## Designing so this hurts less The operational lesson interviewers are usually fishing for is that **DNS-based failover is not instant**, and its worst-case delay is not set by the record's lifetime alone. It is the sum of every layer that holds a copy plus every connection that never needed to re-resolve. A record change propagates to new connections from processes that re-resolve; it does not migrate live traffic. That is why work that must move quickly does not rely on a name change alone: connections are drained or capped in lifetime so clients are forced to re-resolve, client-side caches are bounded rather than left at a runtime's default, and the number of caching layers between an application and the authoritative data is kept small enough to reason about. When a fast cutover really is required, moving the address — or steering at a load balancer that owns the connections — is more reliable than hoping every layer expires on time.

  • Why is an already-established connection the case people forget?
    Because it looks like a caching problem but involves no cache. A name is resolved only when a new connection is opened; an existing or pooled connection already has a peer address and never asks again. Flushing every cache on the host changes nothing until that connection is closed and a new one is created.
  • A fresh lookup returns the new address, so is it worth flushing the local resolver cache?
    No — the fresh lookup already went through that cache and returned the new value, so the cache is not holding the old one. Flushing is a reasonable step when the host itself still returns the stale address, but here it only wastes time and hides the fact that the stale copy is inside the process or in an open connection.
  • What does this imply about using a DNS record change to fail over quickly?
    That the cutover is best-effort and its true delay is the slowest layer holding a copy, plus every long-lived connection that never re-resolves. Plan for connection lifetimes to be bounded and client caches to be configured explicitly, or move traffic at a component that owns the connections rather than relying on the name change alone.

saying these in an interview costs you the question

  • Claims Linux has no DNS cache at all
  • Flushes caches when a fresh lookup already returns the new address
  • Forgets that open connections never re-resolve
  • Assumes every client honours the record's lifetime
  • Restarts the resolver instead of the process holding the address

context