skip to content

Workloads in a cluster intermittently fail name resolution: most lookups succeed, but a small fraction time out after roughly five seconds. How would you investigate, and what causes are plausible?

level: seniorimportance: should knowfreq 38%

answer

  1. ~5s = resolver timeout → packets dropped, not slow
  2. scope: which names / nodes / clients / when
  3. NXDOMAIN share = ndots amplification
  4. conntrack insert-failed = A+AAAA race
  5. node-local cache + scale/protect the DNS pods

basics

~20 s

Roughly five seconds is the C library resolver's retry timeout, so a query is being dropped rather than answered slowly. Look for dropped UDP queries from parallel A and AAAA lookups colliding in connection tracking, and for overloaded or CPU-throttled DNS pods. Confirm with DNS server metrics, then deploy a node-local cache.

solid answer

~60 s

The five-second figure is the tell: it is the resolver's default timeout before retrying, so queries are being **lost**, not answered slowly. Triage in layers. 1. **Scope it.** All namespaces or one? All nodes or a few? Cluster names, external names, or both? Failures only for external names point at ndots amplification plus upstream latency; failures on specific nodes point at node-local networking. 2. **Server side.** Check CoreDNS: replica count against cluster size, CPU throttling against its limits, restarts or evictions, request-duration and response-code metrics. A high share of NXDOMAIN indicates search-list amplification inflating load. 3. **Client side.** Read `/etc/resolv.conf` in an affected pod: `ndots:5` plus external calls multiplies queries and thus multiplies exposure. 4. **Datapath.** Parallel A and AAAA queries leaving one socket can collide when connection-tracking entries are created for the translation to a DNS pod, dropping one; `conntrack -S` insert failures corroborate. Fixes: node-local DNS cache (also forwarding upstream over TCP), scale and protect CoreDNS with requests and a disruption budget, lower `ndots` for egress-heavy workloads.

code

bash · 13 lines
bash
kubectl -n kube-system get deploy coredns
kubectl -n kube-system top pod -l k8s-app=kube-dns
kubectl -n kube-system get events --field-selector reason=Evicted

# query a specific DNS replica instead of the Service address
kubectl -n kube-system get pods -l k8s-app=kube-dns -o wide
dig @10.244.1.12 payments.billing.svc.cluster.local

# on an affected node
conntrack -S | grep insert_failed

# reproduce from a client pod
for i in $(seq 1 200); do /usr/bin/time -f %e dig +short api.example.com. >/dev/null; done

go deeper

for a junior

Recognise that resolution can fail intermittently and know to check whether the cluster DNS pods are healthy and how many replicas are running.

for a middle

Explain the meaning of the five-second timeout, check both server metrics and the client's resolv.conf, and name the node-local cache as a fix.

for a senior

Run the layered narrowing with evidence at each step, distinguish server saturation from the datapath race using metrics and connection-tracking counters, and pair the immediate fix with capacity and protection work.

for a principal

Treat cluster DNS as a tier-0 dependency: define an availability and latency objective for it, budget capacity against query amplification, and design its rollout and disruption behaviour so upgrades never remove a large fraction of serving capacity.

## Read the number first A lookup that fails at about five seconds is not a slow server. The standard C library resolver waits a fixed timeout — five seconds by default — then retries or gives up. So the query, or its answer, never arrived. That single observation eliminates "the upstream resolver is slow" as the primary hypothesis and points at packet loss or a saturated server dropping requests. ## Step 1 — scope the failure Gather cheaply before theorising: - **Which names?** Only external names failing suggests the search-list walk and upstream forwarding are involved. Cluster names failing too suggests the DNS pods themselves. - **Which nodes?** Concentration on a few nodes points at node-local causes: a saturated connection-tracking table, an unhealthy DNS pod scheduled there, or an overloaded node. - **Which clients?** A single application failing while others are fine often means that application does the most lookups — no caching, new resolution per request. - **When?** Correlate with deploys, node scale-ups, or a rollout of the DNS Deployment. ## Step 2 — the server side CoreDNS is a Deployment with a small default replica count that many clusters never revisit as they grow. - **Capacity.** Are two replicas serving thousands of pods? Check whether CPU throttling is occurring against the container's limit — a throttled DNS server drops queries in bursts, which produces exactly this intermittent pattern. - **Health.** Restarts, OOM kills, or evictions. If the DNS pods have no resource requests and no priority class, they compete with application pods and can be evicted under pressure. - **Metrics.** Request duration, and responses broken down by response code. A high NXDOMAIN share is the fingerprint of search-list amplification: the server is spending its capacity answering the expansion attempts for external names. - **Upstream.** Cache misses for external names are forwarded; if the upstream resolver is slow or rate-limiting, forwarded queries queue and time out. Compare in-cluster name latency against external name latency to separate the two. - **Isolate a replica.** Query a specific DNS pod's address directly rather than the Service address. If one replica misbehaves and others are fine, you have found it; if all are fine when queried directly, suspect the path to them. ## Step 3 — the client side Read the effective `/etc/resolv.conf` in an affected pod. With `ndots:5`, every external hostname triggers several suffix attempts, each usually issued as a parallel A and AAAA pair. That multiplies both server load and the number of chances any single query is lost. An application that resolves per request rather than reusing connections multiplies it again. ## Step 4 — the datapath race The well-known cause of precisely-five-second stalls: a resolver sends the A and AAAA queries from the **same source socket** at nearly the same instant, to the DNS Service address. Both packets need a connection-tracking entry for the translation to a backing DNS pod, and concurrent creation of two entries for the same tuple can fail, so one packet is dropped. The application then waits the full resolver timeout. Corroborate with `conntrack -S` on affected nodes — insert-failed counters climbing in step with the incidents — and by checking whether the connection-tracking table is near capacity, which produces its own drops. Workarounds, weakest to strongest: resolver options that stop issuing the two queries simultaneously on one socket; disabling AAAA lookups where the workload is v4-only; and the structural fix, a node-local cache that answers from the node and forwards upstream over TCP, taking the racing UDP pattern out of the picture. ## Step 5 — the fixes, in order of leverage 1. **Node-local DNS cache.** A per-node cache pod that pods query at a fixed local address. It removes most cross-node DNS traffic, absorbs the repeated NXDOMAIN answers, and uses TCP upstream, which eliminates the UDP race. It is the standard remedy at scale precisely because it requires no application change. 2. **Right-size and protect the DNS Deployment.** Scale replicas with cluster size (an autoscaler proportional to nodes or cores is common), set requests so it is not throttled, give it a critical priority class, add anti-affinity so replicas do not share a node, and a disruption budget so maintenance cannot take them all. 3. **Reduce query volume.** Lower `ndots` for egress-heavy workloads, fully qualify hot external names, and fix applications that resolve on every request instead of reusing connections. Raising the cache TTL for external names helps too. 4. **Watch the rollout path.** Rolling the DNS Deployment briefly removes capacity; with only two replicas that is a large fraction. Ensure readiness gating and surge behaviour keep enough serving capacity during upgrades. ## How to present this in an interview Lead with the meaning of the five seconds, then show the layered narrowing — names, nodes, clients, time — then name both plausible causes (server saturation and the query race) and say how you would distinguish them with evidence. Finish with the node-local cache as the change with the best leverage, and the capacity and protection work that stops it recurring.

  • Why does a node-local DNS cache reduce five-second stalls even when the DNS server itself is healthy?
    Pods query a cache on their own node, so most lookups never traverse the pod network or get translated to a DNS pod address, removing the connection-tracking step where the parallel A and AAAA queries collide. Repeats — including the NXDOMAIN answers produced by search-list expansion — are served locally, cutting load sharply. Forwarding upstream over TCP removes the racing UDP pattern for the remainder.
  • What in the DNS server's metrics would tell you the problem is client-side query amplification rather than server capacity?
    A large share of responses with the NXDOMAIN code relative to successful answers, combined with request rates far above what the application's real lookup volume should be. That pattern means the server is mostly answering search-suffix attempts for names that will ultimately resolve elsewhere. Capacity problems instead show as rising request duration, CPU throttling and restarts under a request rate that is not obviously inflated.

saying these in an interview costs you the question

  • Treating a five-second stall as a slow upstream resolver rather than a dropped query
  • Scaling DNS replicas as the first and only action, without measuring where queries are lost
  • Ignoring that CPU limits can throttle the DNS pods into dropping requests
  • Never checking the client's resolv.conf, so ndots amplification goes unnoticed
  • Concluding DNS is fine because a manual dig succeeded once

context