How does an individual Pod in a Kubernetes StatefulSet get a stable DNS name, and what role does a headless Service play in that?
answer
- clusterIP: None → no VIP, DNS returns Pod IPs
- spec.serviceName = governing Service
- pod.service.namespace.svc.cluster.local
- hostname = pod name, subdomain = service
- name stable, IP is not; publishNotReadyAddresses for bootstrap
basics
~10 sThe StatefulSet names a governing Service in spec.serviceName. That Service is headless (clusterIP: None), so DNS returns Pod addresses rather than a virtual IP, and each Pod resolves at <pod-name>.<service>.<namespace>.svc.cluster.local.
solid answer
~50 sA StatefulSet's `spec.serviceName` names the **governing Service**, which you create yourself as headless by setting `clusterIP: None`. Headless means the Service gets no virtual IP and no kube-proxy load balancing; instead cluster DNS returns the addresses of the backing Pods directly. The StatefulSet controller sets each Pod's `hostname` to the Pod name and its `subdomain` to the governing Service name. Cluster DNS therefore publishes a per-Pod record: `<pod-name>.<service-name>.<namespace>.svc.cluster.local` — for example `kafka-1.kafka-hs.data.svc.cluster.local`. That name is stable for the ordinal. The Pod's IP changes when it is rescheduled; the name resolves to the new IP. Peers can therefore be configured with a fixed member list at deploy time, which is exactly what quorum systems need. Querying the headless Service name itself returns all ready Pod IPs, giving cheap peer discovery. Set `publishNotReadyAddresses: true` when members must find each other *before* they pass readiness, as during cluster bootstrap.
code
yaml · 13 linesapiVersion: v1
kind: Service
metadata:
name: kafka-hs
namespace: data
spec:
clusterIP: None
publishNotReadyAddresses: true
selector:
app: kafka
ports:
- name: peer
port: 9093go deeper
Recall that clusterIP: None makes a Service headless and that each StatefulSet Pod then gets its own DNS name built from the Pod name and the Service name.
Give the full FQDN pattern, explain that the controller sets hostname and subdomain, and state clearly that names are stable while IPs are not.
Cover readiness gating and publishNotReadyAddresses, the two-Service pattern for peer versus client traffic, SRV-based discovery, and client-side DNS caching pitfalls.
Discuss membership and discovery strategy overall — static member lists versus SRV discovery versus an operator managing membership — and how identity choices constrain cross-cluster or multi-region topologies.
## Why a normal Service is not enough A standard ClusterIP Service allocates one virtual IP and load-balances connections across all Ready endpoints. That is ideal for stateless traffic and useless for a clustered database: if `kafka-0` needs to replicate to `kafka-1` specifically, resolving a name that lands on a random broker is wrong. Clustered systems need to address *members*, not the set. ## The headless Service Setting `spec.clusterIP: None` makes a Service headless. Concretely: - No virtual IP is allocated, and kube-proxy programs no forwarding rules for it. - Cluster DNS answers a query for the Service name with the IP addresses of the backing Pods themselves (multiple A/AAAA records), instead of a single virtual IP. - Because the endpoints are real Pod IPs, clients connect directly to Pods. On its own that already gives you discovery: one lookup, all the members. ## Per-Pod records The StatefulSet controller adds identity on top. It sets each Pod's `spec.hostname` to the Pod name (`kafka-1`) and `spec.subdomain` to the governing Service named in `spec.serviceName` (`kafka-hs`). Cluster DNS uses those fields to publish an individual record per Pod: ``` <pod-name>.<governing-service>.<namespace>.svc.cluster.local kafka-1.kafka-hs.data.svc.cluster.local ``` Inside the same namespace the short form `kafka-1.kafka-hs` resolves too, thanks to the Pod's DNS search path. The container's own `hostname` command also returns `kafka-1`, and its FQDN matches — which matters for systems like Kafka, Cassandra and ZooKeeper that advertise their own hostname to peers. The governing Service must actually exist, and the StatefulSet does not create it for you. A missing or misnamed `serviceName` is a classic bug: Pods start fine, but peer resolution silently fails and the cluster never forms. ## What is and is not stable The **name** is stable; the **IP is not**. When `kafka-1` is rescheduled onto a different node it gets a new Pod IP, and DNS is updated to match. This is the whole point: peers configured with `kafka-1.kafka-hs` keep working across node failures, rolling updates and cluster upgrades. Any application that caches a resolved IP indefinitely defeats it — JVM services in particular should have their DNS cache TTL set to something short rather than the historical forever. ## Readiness and bootstrap By default a headless Service publishes only *ready* endpoints, and per-Pod DNS records follow the same rule. That creates a bootstrap deadlock for systems where a node is not ready until it has joined a quorum, but cannot join a quorum until it can resolve its peers. The fix is `publishNotReadyAddresses: true` on the headless Service, which publishes records for Pods regardless of readiness so members can discover each other while still starting up. A common production layout uses **two Services**: a headless governing Service with `publishNotReadyAddresses: true` for peer-to-peer traffic and identity, and a separate normal ClusterIP Service for client traffic, which correctly load-balances over ready members only. Some systems add a third selector-labelled Service pointing only at the current primary. ## SRV records and discovery Headless Services also publish SRV records for their named ports, in the form `_<port-name>._<protocol>.<service>.<namespace>.svc.cluster.local`, resolving to the per-Pod names. Systems that discover peers dynamically use an SRV lookup rather than a hardcoded member list, so the set can grow without editing config. ## Practical checks When peer discovery breaks, verify in order: does the headless Service exist with exactly the name in `spec.serviceName`; does its selector actually match the Pod labels; are endpoints populated; does a DNS lookup from inside a Pod resolve both the Service name and an individual Pod name; and is readiness gating the records you expect. Nearly every StatefulSet DNS incident is one of those five.
- A StatefulSet's Pods start normally but cannot resolve each other. What do you check first?Whether the governing Service named in `spec.serviceName` actually exists in the same namespace, is headless, and has a selector matching the Pod labels — the StatefulSet does not create it and does not complain when it is missing. Then confirm the endpoints are populated and run an in-cluster nslookup of both the Service name and a single Pod name.
- Why would you set publishNotReadyAddresses on the headless Service?Because DNS normally publishes only ready endpoints, and many clustered systems cannot become ready until they have contacted their peers — a deadlock at first bootstrap. Publishing not-ready addresses lets members resolve each other while still starting. Client traffic should go through a separate normal Service that keeps the ready-only behaviour.
saying these in an interview costs you the question
- Believing the StatefulSet creates its governing Service automatically.
- Saying Pod IPs are stable in a StatefulSet — only the DNS names are.
- Expecting per-Pod DNS records without a headless Service, or from a regular ClusterIP Service.
- Pointing client traffic at the headless Service and assuming it load-balances like a ClusterIP.
- Forgetting that unready Pods are absent from DNS by default, which breaks first-time cluster bootstrap.