skip to content

A Kubernetes pod's /etc/resolv.conf contains a search list and the line 'options ndots:5'. Explain what each does, and why together they can make lookups of external names slow and generate large amounts of DNS traffic.

level: middleimportance: must knowfreq 44%

answer

  1. search = suffixes appended to short names
  2. ndots:5 = fewer than 5 dots → try suffixes first
  3. api.example.com = 2 dots → 4 NXDOMAIN, doubled by A+AAAA
  4. fixes: trailing dot, dnsConfig ndots, node-local cache
  5. NXDOMAIN share in metrics = the tell

basics

~20 s

The search list is appended to short names so bare Service names resolve. ndots:5 means any name with fewer than five dots is tried against every search suffix first. An external name like api.example.com has two dots, so it produces several failed cluster lookups before the real one.

solid answer

~50 s

`search` supplies the suffixes appended to unqualified names — typically `<ns>.svc.cluster.local`, `svc.cluster.local`, `cluster.local` — which is what makes a bare Service name resolve. `options ndots:5` sets the threshold for treating a name as already absolute. With `ndots:5`, any name containing **fewer than five dots** is first tried with each search suffix appended, and only then as given. Kubernetes uses 5 because the longest cluster form, `_port._tcp.svc.ns.svc.cluster.local`, needs it. The cost lands on external names. `api.example.com` has two dots, so it walks the whole search list — several NXDOMAIN round trips, each usually doubled by a parallel A and AAAA query — before the correct answer. That is 6–8 queries where 2 would do: added tail latency on every cold lookup and multiplied load on CoreDNS. Mitigations: use a trailing dot to make names absolute, override `ndots` per pod via `dnsConfig`, and deploy a node-local DNS cache.

code

bash · 8 lines
bash
cat /etc/resolv.conf
# search billing.svc.cluster.local svc.cluster.local cluster.local
# options ndots:5

dig +search +trace api.example.com | grep -c 'NXDOMAIN'

# absolute form: no suffix expansion at all
dig api.example.com.

go deeper

for a junior

Explain that the search list makes short Service names work and that ndots:5 makes short external names get tried against those suffixes first, which wastes lookups.

for a middle

Walk a concrete external name through the full expansion, count the queries including the A/AAAA doubling, and name at least two fixes with their tradeoffs.

for a senior

Connect it to observed symptoms — external-call latency floor, high NXDOMAIN share, five-second stalls — and argue for the node-local cache as the change that needs no application edits.

for a principal

Treat it as a platform default with a cost: decide whether the cluster optimises for in-cluster or egress-heavy workloads, whether ndots is set by admission defaults per workload class, and what DNS capacity headroom the amplification implies.

## The two directives `/etc/resolv.conf` in a pod looks roughly like: ``` nameserver 10.96.0.10 search billing.svc.cluster.local svc.cluster.local cluster.local options ndots:5 ``` **`search`** lists suffixes the resolver appends to a name that is not fully qualified. It is what allows `payments` to mean `payments.billing.svc.cluster.local` inside the `billing` namespace, and `payments.shop` to mean the same Service in another namespace. **`ndots:N`** decides *when* the resolver bothers with the search list at all. The rule in the standard C library resolver: if the name contains **N or more dots**, try it as an absolute name first; otherwise, try every search suffix first and only fall back to the name as given. The default outside Kubernetes is `ndots:1`, meaning almost any dotted name is treated as absolute. ## Why Kubernetes chose 5 The cluster's own names need generous expansion. The longest supported short form is an SRV query like `_http._tcp.payments.billing`, which has four dots and must still be expanded against `svc.cluster.local` to resolve. Setting the threshold at 5 guarantees that every legitimate abbreviated cluster name is expanded. It optimises for in-cluster names — and that choice is exactly what taxes external ones. ## Walking a real lookup A pod in namespace `billing` resolves `api.example.com` (2 dots, under the threshold of 5). The resolver tries, in order: 1. `api.example.com.billing.svc.cluster.local` → NXDOMAIN 2. `api.example.com.svc.cluster.local` → NXDOMAIN 3. `api.example.com.cluster.local` → NXDOMAIN 4. (any additional suffix inherited from the node) → NXDOMAIN 5. `api.example.com` → answer And because modern resolvers request A and AAAA in parallel, roughly double those packets. Eight round trips to CoreDNS to answer one question. ## What that actually costs **Latency.** Each failed attempt is a round trip to the DNS Service and back. On a healthy cluster that is sub-millisecond and invisible; on a loaded or throttled CoreDNS it becomes visible tail latency on every cold connection, and it happens repeatedly in applications that do not cache resolutions. **Load.** Cluster DNS query volume is multiplied several-fold. A cluster whose workloads talk mostly to external APIs can spend the majority of its DNS capacity generating NXDOMAIN answers. This is a very common root cause behind "CoreDNS is at its CPU limit" investigations. **Amplified exposure to a known race.** Parallel A and AAAA queries sent from the same socket to the same destination can, on some kernels, collide when connection-tracking entries are created for the two UDP flows, causing one to be dropped. The application then waits for the resolver's default timeout — the classic **five-second** stall — before retrying. Every additional search-suffix attempt is another chance to hit it, so ndots amplification and the 5-second symptom appear together. ## Mitigations, roughly in order of preference 1. **Fully qualify external names, with a trailing dot.** `api.example.com.` is absolute: no suffix is ever appended. This is free and precise, but it requires touching application configuration and some libraries or URL parsers mishandle the trailing dot — validate before rolling it out broadly, particularly where TLS certificate matching is involved. 2. **Lower `ndots` per pod via `dnsConfig`.** Setting `options ndots: 1` or `2` on workloads that mostly call outward removes the expansion for their external names. The tradeoff: names *shorter* than the new threshold still expand, so a pod set to `ndots:1` can no longer resolve `payments.shop` as a short form — it needs the fully qualified name. Apply it where in-cluster short names are not used. 3. **Deploy a node-local DNS cache.** A per-node cache answers repeats — including the repeated NXDOMAINs — without crossing the network to a DNS pod. It also typically uses TCP upstream, which sidesteps the UDP conntrack race entirely. This is the standard fix at scale because it needs no application change. 4. **CoreDNS-side expansion (autopath).** A CoreDNS plugin can recognise the search-path walk and answer with the final result directly, collapsing several queries into one. It saves client round trips but requires CoreDNS to track pod namespaces, adding memory and API-watch cost — evaluate rather than enable reflexively. 5. **Cache in the application.** Reusing connections and bounding resolver cache TTL removes most repeat lookups, whatever the cluster does. ## The diagnostic tell If external calls have a latency floor that in-cluster calls do not, or CoreDNS metrics show a large share of NXDOMAIN responses, ndots amplification is the first thing to check. Reproduce with a single lookup traced from a pod and count the queries: the gap between one intended query and the eight observed ones is the entire finding.

  • What breaks if you set ndots to 1 for every pod in the cluster?
    Short in-cluster names stop resolving as expected. With ndots:1, any name containing a dot is treated as absolute, so payments.shop is queried literally instead of expanding to payments.shop.svc.cluster.local and fails. Bare single-label names such as payments still expand, so same-namespace calls survive, but cross-namespace short forms and short SRV forms break. Apply the override to workloads that are known to use fully qualified or external names.
  • Why does a node-local DNS cache help more than tuning ndots alone at large scale?
    It removes the network round trip for repeated lookups, including the repeated NXDOMAIN answers the search walk produces, so cluster DNS load drops sharply without touching any application. It also typically forwards upstream over TCP, which avoids the UDP connection-tracking race that produces five-second resolution stalls. The cost is another per-node component to operate and to reason about during upgrades.
  • How does the trailing dot in 'api.example.com.' change resolution?
    It marks the name as fully qualified, so the resolver skips the search list entirely and queries it directly. That turns roughly eight queries into two. The caveat is library support: some HTTP clients, URL parsers and TLS certificate matchers handle the trailing dot inconsistently, so it should be verified per client rather than applied blindly.

It is like an office mail room told to try every internal department before accepting that a letter is addressed outside the building: correct in the end, but four wasted trips for every external letter.

saying these in an interview costs you the question

  • Thinking ndots controls how many suffixes are tried rather than the dot threshold for skipping them
  • Assuming external DNS is fine because in-cluster resolution feels fast
  • Setting ndots:1 cluster-wide without realising cross-namespace short names break
  • Blaming CoreDNS capacity without measuring the NXDOMAIN share of its traffic
  • Confusing the five-second stall (a resolver timeout after a dropped query) with slow upstream servers

context