skip to content

Cluster DNS and CoreDNS

CoreDNS answers every in-cluster name such as svc.namespace.svc.cluster.local, and its Corefile, the ndots:5 search-domain default and NodeLocal DNSCache decide whether lookups stay fast. Slow or intermittent DNS after a cluster grows is a routine SRE probe.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

6

How does a workload running in a Kubernetes cluster resolve another Service by name? Describe what component answers the query and the DNS name forms available.

level: juniorimportance: must knowfreq 70%

answer

  1. CoreDNS in kube-system, ClusterIP in resolv.conf
  2. <svc>.<ns>.svc.cluster.local → ClusterIP
  3. bare name = own namespace only
  4. SRV: _port._proto.<svc>.<ns>...
  5. ExternalName = CNAME, no proxying

basics

~20 s

CoreDNS runs in the cluster and its Service address is written into every pod's /etc/resolv.conf. A Service resolves as <service>.<namespace>.svc.cluster.local to its ClusterIP. Inside the same namespace the bare name works; across namespaces use <service>.<namespace>.

solid answer

~40 s

Cluster DNS is provided by **CoreDNS**, running as a Deployment in `kube-system` and fronted by a Service (conventionally named `kube-dns`) with a fixed ClusterIP. The kubelet writes that address as the `nameserver` in every pod's `/etc/resolv.conf`, so ordinary resolver calls reach it with no application configuration. The fully qualified form for a Service is `<service>.<namespace>.svc.cluster.local`, and it returns the Service's ClusterIP — one virtual address, not the pod IPs. Shorter forms work because the kubelet also writes a **search** list: within a namespace `payments` is enough, and `payments.billing` reaches another namespace. Ports can be discovered with SRV records at `_<port-name>._<protocol>.<service>.<namespace>.svc.cluster.local`. Pods also get an address-derived name under `pod.cluster.local`, rarely used directly. A `type: ExternalName` Service resolves to a CNAME pointing at an external hostname rather than to any cluster address.

code

bash · 8 lines
bash
cat /etc/resolv.conf
# nameserver 10.96.0.10
# search billing.svc.cluster.local svc.cluster.local cluster.local
# options ndots:5

nslookup payments                     # → payments.billing.svc.cluster.local
nslookup payments.shop                # another namespace
nslookup -type=srv _http._tcp.payments.billing.svc.cluster.local

go deeper

for a junior

Give the fully qualified form, say CoreDNS answers it, and show you know the bare name only works inside your own namespace.

for a middle

Add the search-list mechanism, SRV records for ports, and the distinction between a ClusterIP answer and pod addresses.

for a senior

Emphasise the diagnostic split — resolution succeeding while connections fail means look at endpoints, not DNS — and mention ExternalName's CNAME-only nature and its TLS pitfalls.

for a principal

Frame DNS as a naming contract that decouples deployment topology from configuration, and note the operational consequence: it becomes a hard dependency on every request path and must be treated as tier-0 infrastructure.

## Why cluster DNS exists Pod IPs are ephemeral — a rescheduled or scaled workload has different addresses every time — so nothing can be configured with them. A Service supplies a stable virtual address, and DNS supplies a stable *name* for that address. Together they mean an application ships with a hostname in its configuration and never learns anything about placement. ## The component **CoreDNS** is the cluster DNS server: a small plugin-chained DNS server whose `kubernetes` plugin watches Services and EndpointSlices through the API server and answers queries from that live view. It runs as a Deployment (usually two replicas by default) in the `kube-system` namespace, exposed by a Service whose ClusterIP is fixed at cluster creation — conventionally the tenth address of the service CIDR, and conventionally still named `kube-dns` for compatibility with older tooling. The kubelet writes that ClusterIP into every pod's `/etc/resolv.conf` as the `nameserver`. Because the standard C library resolver reads that file, every language's name lookup automatically uses cluster DNS without any library or sidecar. ## The name forms **Services (A/AAAA).** ``` <service>.<namespace>.svc.<cluster-domain> ``` with `cluster.local` as the near-universal cluster domain. For a normal Service this returns exactly one address: the **ClusterIP**. It does not return pod addresses, and it does not load balance — the spread across pods happens after the packet is sent, when the ClusterIP is translated to a backing pod. **Shorter forms.** `/etc/resolv.conf` also carries a `search` list, typically: ``` search <namespace>.svc.cluster.local svc.cluster.local cluster.local ``` The resolver appends each suffix in turn to an unqualified name. So from a pod in `billing`: - `payments` → `payments.billing.svc.cluster.local` (first suffix matches) - `payments.shop` → `payments.shop.svc.cluster.local` (second suffix) - `payments.shop.svc.cluster.local` → resolved as given The common beginner error is expecting a bare name to reach another namespace. It cannot: the first suffix pins it to the caller's own namespace. **Ports (SRV).** ``` _<port-name>._<protocol>.<service>.<namespace>.svc.cluster.local ``` This returns the port number together with the target name, so a client can discover the port rather than hard-coding it — useful when ports are assigned per environment. It requires the Service port to have a `name`. **Pods.** A record of the form `<pod-ip-with-dashes>.<namespace>.pod.cluster.local` resolves to that address. It is rarely used directly by applications; it exists so that constructed names resolve consistently. **ExternalName Services.** `spec.type: ExternalName` with `spec.externalName: db.example.com` creates no ClusterIP and no proxying. The DNS answer is a **CNAME** to the external name, so a workload keeps addressing an in-cluster name while the traffic goes elsewhere. It is a DNS-level indirection only: nothing rewrites TLS SNI or HTTP Host headers, which is why it surprises people with certificate mismatches. **Headless Services.** A Service with `clusterIP: None` has no virtual address, and its name resolves instead to the set of ready backing pod addresses — a different behaviour worth knowing exists, used for peer discovery and client-side load balancing. ## What to verify when a name fails From a debug pod: `cat /etc/resolv.conf` to see the nameserver and search list; `nslookup payments.billing.svc.cluster.local`; then compare with `kubectl get svc -n billing payments`. Three failures dominate: the Service does not exist in the namespace you assumed; the name was left unqualified and hit the caller's namespace instead; or the Service exists but has no ready endpoints, in which case the ClusterIP still resolves and connections fail afterwards — which is an important distinction, because a successful lookup does not imply a working backend. ## The mental model to carry DNS in Kubernetes answers "what address", never "which pod is healthy". Health and distribution are the Service's job. A name resolving correctly while requests fail is normal and points downstream, not at DNS.

  • A pod in namespace 'billing' cannot reach a Service named 'payments' that lives in namespace 'shop', using just the hostname 'payments'. Why?
    The search list puts the caller's own namespace first, so the unqualified name expands to payments.billing.svc.cluster.local and returns NXDOMAIN because no such Service exists in billing. Cross-namespace addressing requires at least payments.shop, or the fully qualified payments.shop.svc.cluster.local. Nothing about namespaces blocks the traffic itself — this is purely name expansion.
  • A Service name resolves successfully but every connection to it is refused or times out. Does that rule DNS in or out?
    It rules DNS out. A normal Service's name resolves to its ClusterIP as long as the Service object exists, entirely independently of whether any pod is behind it. So a good lookup with failing connections points to the Service having no ready endpoints, to the wrong target port, or to the application itself — check the endpoints and pod readiness next.

saying these in an interview costs you the question

  • Expecting a bare Service name to resolve across namespaces
  • Believing a Service's A record returns pod IPs and that DNS does the load balancing
  • Thinking applications must be configured with the DNS server address themselves
  • Assuming a successful lookup proves there are healthy backends
  • Treating ExternalName as a proxy that rewrites Host headers or TLS SNI

context

open as a page

A Kubernetes pod's /etc/resolv.conf contains a search list and the line 'options ndots:5'. Explain what each does, and why together they can make lookups of external names slow and generate large amounts of DNS traffic.

level: middleimportance: must knowfreq 44%

basics

~20 s

The search list is appended to short names so bare Service names resolve. ndots:5 means any name with fewer than five dots is tried against every search suffix first. An external name like api.example.com has two dots, so it produces several failed cluster lookups before the real one.

open as a page

What DNS records does a Kubernetes Service created with clusterIP: None produce, and how do StatefulSet pods obtain stable individual DNS names from it?

level: middleimportance: should knowfreq 45%

basics

~20 s

With clusterIP: None there is no virtual address: the Service name resolves to the addresses of all ready backing pods. Paired with a StatefulSet through its serviceName, each pod also gets its own stable name, <pod>.<service>.<namespace>.svc.cluster.local.

open as a page

What do the Kubernetes pod spec fields dnsPolicy and dnsConfig control, and what are the available dnsPolicy values?

level: middleimportance: should knowfreq 34%

basics

~20 s

dnsPolicy chooses which resolver configuration a pod gets: ClusterFirst (the default, cluster DNS), ClusterFirstWithHostNet (cluster DNS for pods sharing the node network namespace), Default (inherit the node's resolver settings), or None (supply everything yourself). dnsConfig adds or replaces nameservers, search domains and resolver options.

open as a page

Workloads in a cluster intermittently fail name resolution: most lookups succeed, but a small fraction time out after roughly five seconds. How would you investigate, and what causes are plausible?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Roughly five seconds is the C library resolver's retry timeout, so a query is being dropped rather than answered slowly. Look for dropped UDP queries from parallel A and AAAA lookups colliding in connection tracking, and for overloaded or CPU-throttled DNS pods. Confirm with DNS server metrics, then deploy a node-local cache.

open as a page

You own cluster DNS for a large, busy multi-tenant Kubernetes cluster. How do you make name resolution scale and stay reliable, and how do you decide how much capacity it needs?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Treat cluster DNS as tier-0: model query volume from pods times lookups per request times search-list amplification, add a node-local cache to absorb repeats, scale and protect the DNS Deployment with requests, anti-affinity, a critical priority class and a disruption budget, and constrain tenants from bypassing or inflating it.

open as a page