skip to content

What DNS records does a Kubernetes Service created with clusterIP: None produce, and how do StatefulSet pods obtain stable individual DNS names from it?

level: middleimportance: should knowfreq 45%

answer

  1. clusterIP: None → A records for ready pods
  2. StatefulSet serviceName → <pod>.<svc>.<ns>.svc.cluster.local
  3. unready excluded → publishNotReadyAddresses for quorum bootstrap
  4. gRPC / peer discovery use case
  5. client DNS cache TTL decides correctness

basics

~20 s

With clusterIP: None there is no virtual address: the Service name resolves to the addresses of all ready backing pods. Paired with a StatefulSet through its serviceName, each pod also gets its own stable name, <pod>.<service>.<namespace>.svc.cluster.local.

solid answer

~40 s

Setting `clusterIP: None` makes a Service **headless**: no virtual IP is allocated and no proxying happens. Cluster DNS answers the Service name with A/AAAA records for every **ready** backing pod, so the client sees the real member addresses and decides for itself how to connect. A StatefulSet that names such a Service in `spec.serviceName` additionally gets per-pod records: ``` <pod-name>.<service>.<namespace>.svc.cluster.local ``` Since StatefulSet pod names are stable and ordinal (`kafka-0`, `kafka-1`), those names survive rescheduling even though the addresses change. That is how clustered systems configure their seed or peer lists. Two gotchas matter. Unready pods are excluded, which deadlocks systems that must form a quorum before they can report ready — `publishNotReadyAddresses: true` fixes that. And clients that cache resolutions indefinitely will keep dialling dead members, so cache TTLs must be sane.

code

yaml · 33 lines
yaml
apiVersion: v1
kind: Service
metadata:
  name: kafka-headless
  namespace: data
spec:
  clusterIP: None
  publishNotReadyAddresses: true   # let members find each other before they are ready
  selector:
    app: kafka
  ports:
    - name: broker
      port: 9092
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: kafka
  namespace: data
spec:
  serviceName: kafka-headless     # source of the per-pod DNS names
  replicas: 3
  selector:
    matchLabels:
      app: kafka
  template:
    metadata:
      labels:
        app: kafka
    spec:
      containers:
        - name: kafka
          image: kafka:3.7

go deeper

for a junior

Say that clusterIP: None means the name returns the pod addresses instead of one virtual address, and that StatefulSet pods get individual names.

for a middle

Give the exact per-pod name form, tie it to serviceName, and explain a real use such as peer discovery or gRPC client-side balancing.

for a senior

Lead with the operational edges — readiness exclusion and publishNotReadyAddresses, TTL and connection-pool caching after rollouts — and when a headless Service is the wrong choice.

for a principal

Discuss where the load-balancing decision should live: pushing membership to clients buys per-request balancing and direct member addressing, but makes every client responsible for health checking, re-resolution and failure handling.

## What headless means A normal Service allocates a ClusterIP: a virtual address that never lives on an interface and is translated to a backing pod by the node's datapath. Setting `spec.clusterIP: None` disables all of that. There is no virtual address, no translation, and no proxying. The Service becomes purely a **grouping and naming** construct over a set of pods. Cluster DNS then changes behaviour. Instead of one A record for a virtual address, it answers the Service name with **one address record per ready backing pod**. A client resolving `cassandra.data.svc.cluster.local` receives the full membership list and connects directly to whichever members it wants. ## Why anyone wants this Three reasons dominate: 1. **Peer discovery.** Distributed databases and queues — Cassandra, etcd, Kafka, Elasticsearch, Zookeeper — need each node to know its peers. Resolving the headless name is a bootstrap mechanism that needs no external registry. 2. **Client-side load balancing.** Long-lived multiplexed connections, notably gRPC over HTTP/2, do not benefit from connection-level distribution: one connection carries everything and pins to one backend. Given all member addresses, a gRPC client can open a subchannel per address and balance per request. 3. **Addressing a specific member.** Sharded and primary/replica systems need to talk to *this* instance, not any instance — impossible through a single virtual address. ## Per-pod names with StatefulSets A StatefulSet declares `spec.serviceName: <headless-service>` — the "governing service". This gives each pod its own DNS name: ``` <pod-name>.<service-name>.<namespace>.svc.cluster.local ``` for example `kafka-1.kafka-headless.data.svc.cluster.local`. Because a StatefulSet assigns stable ordinal names, and re-creates a replacement under the same name, that DNS name is durable across reschedules even though the address behind it changes. This is the whole point: the identity is the name, and configuration files, seed lists and peer certificates can reference it. SRV records complete the picture: querying `_<port-name>._tcp.<service>.<namespace>.svc.cluster.local` returns one entry per member with its port and target name, which is how some clients enumerate a cluster in a single query. ## Readiness and the bootstrap deadlock By default only **ready** endpoints appear. That is usually right — you do not want clients steered at a member that cannot serve — but it creates a chicken-and-egg failure on first start: a quorum-based system's members are unready until they find each other, and they cannot find each other because unready members are not published. The result is a cluster that never forms. The fix is `publishNotReadyAddresses: true` on the headless Service, which publishes endpoints regardless of readiness. Use it deliberately, and only where the client tolerates dialling members that are still starting — which peer protocols generally do, since they retry. ## Caching pitfalls With a virtual address, a stale DNS answer is harmless: the address is stable and the datapath handles membership. With a headless Service, the DNS answer *is* the membership, so caching directly determines correctness. - CoreDNS serves these records with a short TTL, but a client that caches beyond it keeps dialling departed pods. - The classic offender is a JVM caching resolutions for the process lifetime under a restrictive security-manager-era default; long-running services must set a bounded cache TTL. - Connection pools that resolve once at startup have the same problem in a different layer: the pool holds addresses of pods that no longer exist and never re-resolves. So when a client keeps hitting a deleted replica after a rollout, look at cache TTL and pool re-resolution before suspecting the cluster. ## Other properties worth stating - **No ordering guarantee.** DNS returns the member set, not a ranking; do not treat the first answer as a primary. - **Selectorless headless Services** (no `selector`) let you manage EndpointSlices by hand, which is how an external system is given an in-cluster name that resolves to its real addresses. - **Not the same as ExternalName.** An ExternalName Service answers with a CNAME to an outside hostname and involves no pods at all. - **Policy still applies.** Discovering an address does not mean you may reach it; enforcement is separate. ## Answering well Say what the record set is, name the per-pod form and where it comes from, then show the two operational edges — readiness exclusion versus `publishNotReadyAddresses`, and client caching turning a membership answer stale. Those edges are what the interviewer is checking for.

  • A three-node quorum database deployed as a StatefulSet never finishes forming a cluster on first start. How does the headless Service explain it?
    Its readiness probe only passes once the cluster forms, but by default only ready endpoints are published, so each member resolves the headless name and sees no peers. Nothing can ever become ready. Setting publishNotReadyAddresses: true on the governing Service publishes members regardless of readiness so they can discover each other and complete the bootstrap.
  • Why do gRPC clients often prefer a headless Service over a normal ClusterIP Service?
    gRPC multiplexes many requests over one long-lived HTTP/2 connection, so connection-level distribution behind a virtual address pins all traffic to whichever backend the single connection landed on. Resolving a headless name gives the client every member address, letting it keep a subchannel per backend and balance per request. The tradeoff is that the client now owns re-resolution and health handling.

saying these in an interview costs you the question

  • Thinking a headless Service still has a virtual IP that just is not shown
  • Assuming per-pod DNS names come automatically from any Deployment rather than from a StatefulSet with a governing service
  • Forgetting unready pods are excluded, and being surprised by bootstrap deadlocks
  • Believing the DNS answer order implies a primary or any ranking
  • Ignoring client-side DNS caching, which turns membership answers stale after a rollout

context