A pod cannot resolve the name of another Kubernetes Service. How do you determine whether cluster DNS is at fault, and what role does the search-domain and ndots configuration in the pod's /etc/resolv.conf play?
answer
- resolv.conf: nameserver + search list + ndots:5
- <5 dots → search domains tried first
- Short name works only inside the same namespace
- dnsPolicy Default / hostNetwork → node resolver, no cluster names
- Egress deny-all breaks port 53 first
basics
~20 sTest resolution from inside a pod against the cluster DNS service, using the full name svc.namespace.svc.cluster.local. Pods get a search list and ndots:5 in /etc/resolv.conf, so short names are tried with each search domain appended — which is why cross-namespace short names fail and external lookups make extra queries.
solid answer
~50 sFirst confirm it is really DNS: a name that does not resolve fails differently from a Service with no endpoints (hang) or a wrong port (refused). From inside a pod (or a `netshoot` debug container, since app images often lack tools): ``` cat /etc/resolv.conf nslookup payments.team-a.svc.cluster.local ``` `/etc/resolv.conf` in a pod contains the cluster DNS service IP, a `search` list — `<ns>.svc.cluster.local svc.cluster.local cluster.local` plus any node domains — and `options ndots:5`. `ndots:5` means any name with **fewer than 5 dots** is first tried with each search domain appended before being tried as-is. Consequences: - A bare `payments` resolves inside its own namespace but **not** from another namespace — use `payments.team-a`. - External names like `api.example.com` (2 dots) generate several NXDOMAIN lookups before succeeding, adding latency; a trailing dot (`api.example.com.`) or `dnsConfig` with a lower ndots avoids it. If full names fail too, check CoreDNS pods, their logs, and the kube-dns Service IP.
code
bash · 7 lineskubectl debug -it payments-client-0 --image=nicolaka/netshoot --target=app -- sh
# inside:
cat /etc/resolv.conf
nslookup payments # short name: uses search + ndots
nslookup payments.team-b.svc.cluster.local
dig @10.96.0.10 payments.team-b.svc.cluster.local +short
dig api.example.com. +short # trailing dot skips the search listgo deeper
Know the Service DNS name format <svc>.<ns>.svc.cluster.local and that a short name only works inside the same namespace.
Explain the search list and ndots:5, the extra lookups for external names, and how to test with nslookup or dig from a debug container.
Separate resolver configuration from DNS-server health, spot dnsPolicy and hostNetwork mistakes, and recognise egress NetworkPolicy on port 53 masquerading as a DNS outage.
Consider cluster-wide DNS load: ndots amplification, NodeLocal DNSCache, CoreDNS sizing and autoscaling, and the blast radius of DNS as a shared dependency for every workload.
## What a pod's resolver actually looks like The kubelet writes `/etc/resolv.conf` inside each pod based on the pod's `dnsPolicy`. For the default `ClusterFirst` it looks roughly like: ``` nameserver 10.96.0.10 search team-a.svc.cluster.local svc.cluster.local cluster.local options ndots:5 ``` - `nameserver` is the ClusterIP of the `kube-dns` Service in `kube-system` (the name is historical; the implementation is CoreDNS in modern clusters). - `search` is the list of suffixes tried for unqualified names, starting with the pod's own namespace. - `ndots:5` sets the threshold: a query name containing **fewer than 5 dots** is considered unqualified and is tried against each search suffix *first*, and only then as an absolute name. ## Why ndots:5 matters Two everyday consequences follow directly. **Cross-namespace short names do not work.** From a pod in `team-a`, `payments` becomes `payments.team-a.svc.cluster.local` — correct if the Service lives in `team-a`. If it lives in `team-b`, that lookup NXDOMAINs, then `payments.svc.cluster.local` and `payments.cluster.local` also fail, and finally bare `payments` fails. The fix is to qualify: `payments.team-b` (which expands to `payments.team-b.svc.cluster.local` via the second search entry) or the full FQDN. **External lookups cost extra queries.** `api.example.com` has two dots, below the threshold, so the resolver dutifully asks for `api.example.com.team-a.svc.cluster.local`, `api.example.com.svc.cluster.local`, `api.example.com.cluster.local` — three NXDOMAINs — before the real query. With IPv4+IPv6 (A and AAAA) that is eight queries for one hostname. On a busy service this is measurable latency and a real load source for CoreDNS. Remedies: append a trailing dot to make the name absolute, or set a per-pod `dnsConfig` with `options: [{name: ndots, value: "1"}]`, or use NodeLocal DNSCache to absorb the volume. ## The Kubernetes DNS naming scheme - Service: `<service>.<namespace>.svc.cluster.local` → the Service's ClusterIP. - Headless Service (`clusterIP: None`): the same name returns the **pod IPs** directly, one A record per ready endpoint. - Pod in a StatefulSet with a governing headless Service: `<pod>.<service>.<namespace>.svc.cluster.local` — stable per-pod names. - SRV records exist for named ports: `_http._tcp.<service>.<namespace>.svc.cluster.local`. `cluster.local` is the default cluster domain and can be configured differently at install time, which matters when a manifest hardcodes it. ## Diagnosing, step by step 1. **Get a shell with tools.** Application images increasingly lack `nslookup`, `dig`, even `sh`. Use `kubectl debug -it <pod> --image=nicolaka/netshoot --target=<container>` to attach an ephemeral container in the same network namespace, or run a temporary netshoot pod in the same namespace. 2. **Read `/etc/resolv.conf`.** Wrong nameserver IP, a missing search list, or `dnsPolicy: Default` (which inherits the *node's* resolver and therefore cannot resolve cluster names at all) shows up here immediately. Pods with `hostNetwork: true` need `dnsPolicy: ClusterFirstWithHostNet`, otherwise they inherit the node resolver and cluster DNS silently stops working. 3. **Query the FQDN.** `nslookup payments.team-b.svc.cluster.local` removes search-domain variables. If the FQDN resolves and the short name does not, it is a search/ndots issue, not a DNS-server issue. 4. **Interpret the failure type.** `NXDOMAIN` means the name genuinely does not exist — check the Service name and namespace, and remember that a Service with zero endpoints still resolves, so successful resolution does not imply working traffic. `SERVFAIL` or a timeout points at the DNS server or the path to it. 5. **Query the DNS service directly** to separate resolver config from server health: `dig @10.96.0.10 payments.team-b.svc.cluster.local`. 6. **Check the DNS layer itself.** `kubectl -n kube-system get pods -l k8s-app=kube-dns` for CoreDNS replicas, their restarts and readiness; `kubectl -n kube-system logs -l k8s-app=kube-dns` for errors such as upstream loops or plugin failures; and confirm the `kube-dns` Service has endpoints — CoreDNS reached through a Service with no endpoints produces exactly the timeouts you are chasing. ## Things that look like DNS but are not - **NetworkPolicy blocking egress to port 53.** The moment a namespace gets a default-deny egress policy, name resolution breaks first and everything downstream looks like a DNS outage. Any deny-by-default egress policy must explicitly allow UDP and TCP 53 to the DNS pods. - **Service resolves but connections hang** — that is an endpoints or policy problem, not DNS. DNS answered correctly; the ClusterIP simply has nothing behind it. - **Intermittent resolution failures at scale** — historically caused by conntrack races on UDP and by DNS query amplification from ndots; NodeLocal DNSCache and TCP fallback mitigate both. Holding this order — resolver config, FQDN, failure type, server health — turns "DNS is broken" into a two-minute diagnosis.
- A pod resolves a Service name successfully but connections to it still hang. Is DNS the problem?No. A ClusterIP Service resolves to its virtual IP regardless of whether any pods back it, so successful resolution proves only that the Service object exists. A hang after resolution points at the next layers: an empty or not-ready EndpointSlice so kube-proxy has nothing to DNAT to, a NetworkPolicy silently dropping the packets, or a wrong targetPort. Check kubectl get endpointslices for the Service before spending more time on DNS.
- Why can a NetworkPolicy make every service in a namespace look like a DNS outage?Because as soon as any policy selects a pod for egress, all egress from that pod is denied except what policies explicitly allow — and the first thing almost every application does is a DNS lookup. If the policy does not allow UDP and TCP port 53 to the cluster DNS pods, name resolution fails and every downstream call fails with it, which reads as 'DNS is down' even though CoreDNS is perfectly healthy. Deny-by-default egress policies should always ship with an explicit DNS allow rule.
- What does a pod with hostNetwork: true need in order to resolve cluster Service names?It needs dnsPolicy: ClusterFirstWithHostNet. With hostNetwork enabled and the default ClusterFirst policy, the pod ends up using the node's resolver, which knows nothing about cluster.local, so Service names fail to resolve while external names work. ClusterFirstWithHostNet keeps the cluster DNS server and search list while the pod shares the host network namespace.
The search list is a mail room that automatically appends your department to any unaddressed envelope. Fine while you write to colleagues; the moment you write to another department by first name only, the letter comes back stamped 'no such person here'.
saying these in an interview costs you the question
- Expecting a bare Service name to resolve from a different namespace
- Believing successful DNS resolution means the Service is actually usable
- Diagnosing from the application image without noticing it lacks nslookup or a shell
- Forgetting that hostNetwork pods need dnsPolicy: ClusterFirstWithHostNet
- Overlooking a default-deny egress NetworkPolicy that blocks port 53