Every name lookup on a Linux server pauses for roughly five seconds and then succeeds. The network is otherwise healthy and the returned addresses are correct. What is the host doing, and what would you change?
answer
- the number itself is the clue
- a resolver default, not the network
- something was sent and never answered
- one family's reply may be missing
- timeout and attempts in resolv.conf
basics
~20 sFive seconds is the glibc resolver's default per-query timeout, so the host is waiting out a query that never gets answered before falling back. The usual causes are an unreachable first nameserver or one of the parallel A and AAAA queries being dropped.
solid answer
~50 sThat number is a fingerprint: five seconds is the glibc resolver's default `timeout`, so the pattern says a query went out and nothing came back, and the answer arrived only on the retry. Two causes dominate. First, the **first `nameserver` in `/etc/resolv.conf` is unreachable or silently filtered** — the resolver waits the full timeout, then falls back to the next entry, and it repeats that on every lookup because failures are never remembered. Second, **the parallel A and AAAA pair**: glibc sends both queries back to back from one socket, and a middlebox or resolver that drops or mishandles one leaves the resolver waiting out the timeout for the missing reply. Confirm by timing a lookup through the NSS path against a direct query to each configured server. Fix the cause — remove the dead server, or set `options single-request-reopen`, or `no-aaaa` on a host with no IPv6 — rather than only shortening `timeout`.
code
bash · 2 linestime getent ahosts app.example.com >/dev/null
time dig +short @10.0.0.53 app.example.comgo deeper
Know that resolution can be slow rather than broken, and that /etc/resolv.conf holds a timeout whose default is five seconds. Check whether the first listed nameserver actually answers.
Explain the resolver walk that produces the delay: query the first server, wait the timeout, fall back, and repeat on every lookup since no failure state is kept. Know that A and AAAA queries go out in parallel and that either reply can go missing.
Diagnose by isolating paths rather than tuning blindly: time the NSS path against direct queries to each configured server, separate a loopback stub's behaviour from glibc's, and fix the cause instead of shrinking the timeout.
Frame it as a latency-budget and blast-radius problem: a per-lookup stall multiplies into thread and connection exhaustion across a fleet. Argue for where caching and retry policy should live so failure handling is owned by one observable component.
## Read the number first A distinctive, repeatable delay is a configuration default announcing itself. Five seconds is `RES_TIMEOUT`, the glibc resolver's default per-query wait, changeable with `options timeout:N` in `/etc/resolv.conf`. Roughly five seconds followed by a *correct* answer therefore means: a query was sent, no reply arrived, the resolver waited out its timeout, and a retry succeeded. Nothing is wrong with the data; something is swallowing one query. That framing matters, because the instinctive reading — "DNS is slow" — sends people to the nameserver's own metrics, which will look perfectly healthy since it answered every query it received. ## Cause one: a dead first nameserver The resolver walks the `nameserver` list in order. If the first entry is powered off, firewalled, or listening but silently dropping UDP port 53, there is no error to detect — only silence — so the resolver must wait the timeout before trying the next entry. Because it keeps no state about server health, **every lookup pays that penalty again**. A server that actively refuses is much less painful: an ICMP port-unreachable, or a TCP reset, lets the resolver move on immediately. The pathological case is precisely the one that produces *no* response. The fix is to remove or repair the entry. `options timeout:1 attempts:2` bounds the damage and is a reasonable stopgap, but it is a smaller wound, not a cure, and shortening the timeout too aggressively can cause spurious failures on a genuinely slow-but-working link. ## Cause two: the A and AAAA pair `getaddrinfo()` asks for both address families by default, and glibc sends the A and AAAA queries back to back, in parallel, from the same source port. That is a performance optimisation with a well-known failure mode: some resolvers and middleboxes handle only the first query, or cannot cope with two outstanding queries on one socket, and one of the two replies never arrives. The resolver has an answer for one family and waits out the full timeout for the other, then retries. The distinguishing symptom is that a lookup restricted to a single family returns instantly while the normal dual lookup stalls. `/etc/resolv.conf` offers three relevant switches: - `single-request` — send the two queries sequentially instead of in parallel. - `single-request-reopen` — keep them parallel but use a fresh socket for the second query, which is the usual remedy when a middlebox is confused by two replies to one port. - `no-aaaa` — suppress AAAA queries altogether (glibc 2.36 and later); appropriate only on a host with no IPv6 connectivity at all. ## Isolate before you tune Establish which resolver is actually in front of you. If `/etc/resolv.conf` points at a loopback stub, you are timing that daemon's retry behaviour rather than glibc's, and its timeouts differ; query the upstream servers directly to separate the two. Then compare paths: ```bash time getent ahosts app.example.com >/dev/null time dig +short @10.0.0.53 app.example.com ``` If the direct query to each configured server is fast but the NSS path is slow, the delay is in the resolver's walk, not in any server. If the first configured server is the slow one and the second is fast, you have cause one. If single-family queries are fast and the dual-family path stalls, you have cause two. ## Why it hurts more than it looks A five-second resolution delay rarely stays a five-second delay. Request handlers hold threads or connections while blocked, so a per-lookup stall becomes queue growth, then pool exhaustion, then timeouts several layers up — and the failures surface as symptoms in services that never mention DNS. Meanwhile short-lived processes that resolve once at start-up pay it on every invocation. This is also why a local caching resolver is a genuine architectural improvement rather than a convenience: it collapses repeated lookups into cache hits and puts the retry policy in one component you can observe, instead of leaving every process to rediscover a dead server through a timeout.
- Why does a nameserver that actively refuses connections cause far less delay than one that is simply unplugged?Because a refusal is a signal. An ICMP port-unreachable or a reset tells the resolver immediately that this server will not answer, so it moves to the next entry without waiting. Silence carries no information, so the only way to conclude that nothing is coming is to wait out the full timeout on every lookup.
- Is lowering `options timeout:` to 1 a proper fix?It is a mitigation, not a fix: the broken query still happens, you just wait less for it. It also reduces headroom on a genuinely slow link, which can turn a working-but-late lookup into a failure. Use it to bound damage while you remove the dead nameserver or correct the query behaviour that is losing a reply.
- How would you tell the A/AAAA cause apart from a dead first nameserver?Compare a single-family lookup with the normal dual-family one against the same server. If restricting the query to one family returns instantly while the dual lookup stalls, one of the two parallel replies is being lost. A dead first nameserver instead makes every query slow regardless of family, and the delay disappears when you query the second server directly.
saying these in an interview costs you the question
- Blames the DNS server, whose own metrics look healthy
- Only raises the timeout, leaving the broken query in place
- Never checks whether the first nameserver replies at all
- Assumes a slow lookup means a slow network link
- Ignores that both A and AAAA queries are sent