A Kubernetes StatefulSet's replicas report Ready only after forming a quorum, and they never do. Why, and what does publishNotReadyAddresses on its headless Service change?
answer
- readiness gates the DNS record
- quorum needs peers first
- endpoints all marked ready
- serving keeps the real probe result
- separate peer and client Services
basics
~20 sCluster DNS publishes only ready endpoints, so replicas waiting on a quorum cannot resolve each other: a deadlock. publishNotReadyAddresses makes the EndpointSlice controller mark every endpoint ready, so peer names resolve as soon as pods have IPs.
solid answer
~40 sPer-pod names under a headless Service come from EndpointSlices, and DNS serves only endpoints whose `ready` condition is true. If readiness waits for a quorum, and the quorum needs peers to resolve each other, nothing ever becomes ready. Setting `publishNotReadyAddresses: true` on the Service makes the EndpointSlice controller write every endpoint with `ready: true`, while `serving` still shows the real probe result. Peer names then resolve once a pod has an IP. Because every consumer of that Service now sees unready pods as ready, keep it on a dedicated peer Service named in `spec.serviceName`, and give clients a separate Service without the flag. Also check `podManagementPolicy`: under `OrderedReady`, later replicas are not even created until earlier ones are Ready.
code
yaml · 26 linesapiVersion: v1
kind: Service
metadata:
name: rec-infer-peers
namespace: ranking
spec:
clusterIP: None
publishNotReadyAddresses: true
selector:
app: rec-infer
ports:
- name: peer
port: 7351
---
apiVersion: v1
kind: Service
metadata:
name: rec-infer
namespace: ranking
spec:
selector:
app: rec-infer
ports:
- name: grpc
port: 8471
targetPort: grpcgo deeper
Recall that a headless Service normally lists only ready pods in DNS, and that publishNotReadyAddresses lists unready ones too.
Explain the chain from readiness to the EndpointSlice ready condition to DNS, and why a quorum-gated readiness probe turns that chain into a deadlock.
Show the operational shape: a dedicated peer Service with the flag, a client Service without it, the OrderedReady trap, and terminating members that stay resolvable.
Weigh putting quorum into readiness at all against a readiness that means able to serve, and judge how much of cluster membership should lean on DNS versus the application's own protocol.
## The deadlock Some clustered applications need their members to **find each other before any of them is useful**. Take a recommendation-model inference server, `rec-infer`, run as a 3-replica **StatefulSet**. Replicas elect a shard coordinator by majority vote, so at least 2 of the 3 must be talking. Each replica's **readiness probe** passes only after that vote succeeds and its 4.2 GB model shard is loaded (about 97 seconds on a single-node development cluster on a laptop). The replicas find their peers through the per-pod names of the StatefulSet's **headless Service**, such as `rec-infer-1.rec-infer-peers.ranking.svc.cluster.local`. Now follow the chain: 1. Cluster DNS publishes a per-pod record only for endpoints whose **`ready`** condition is true. 2. The EndpointSlice controller sets `ready` from the pod's readiness. 3. Readiness waits for a quorum. 4. A quorum needs peers to resolve each other's names. Nobody becomes ready, so nobody is published, so nobody becomes ready. The pods run, their logs show lookups that return no such name, and the set never forms. ## What `publishNotReadyAddresses` changes `spec.publishNotReadyAddresses` is a boolean on the **Service**. It tells anything that consumes that Service's endpoints to disregard readiness. Concretely, the EndpointSlice controller writes every endpoint of that Service with **`ready: true`**, whatever the pod's readiness probe says. The `serving` condition still reports the real probe result, so the truth is not lost. Because DNS follows the `ready` condition, each replica's per-pod name resolves as soon as the pod has an IP. The peers find each other, the vote succeeds, and readiness passes. Limits of the fix: - **An IP is still required.** A pod whose sandbox is not set up yet has no IP and gets no endpoint at all. - **Finished pods are dropped.** Pods in a terminal phase (Succeeded or Failed) are not listed. - **Terminating pods stay "ready".** With the flag set, a pod that is shutting down is still written with `ready: true`, so peers keep resolving a departing member. The application's own membership protocol has to cope with that. - **Negative answers get cached.** A replica that looked a peer up before the peer had an IP may have cached the failure. Lookups should retry with backoff. ## Why you want two Services The flag applies to **every consumer** of that Service. If clients used the same Service, they would be sent to replicas that are still loading. So the usual layout is: | Service | `clusterIP` | `publishNotReadyAddresses` | Used by | |---|---|---|---| | `rec-infer-peers` | `None` | `true` | Replicas finding each other; named in the StatefulSet's `spec.serviceName` | | `rec-infer` | assigned | `false` (default) | Ranking clients, which should only reach ready replicas | ```yaml apiVersion: v1 kind: Service metadata: name: rec-infer-peers namespace: ranking spec: clusterIP: None publishNotReadyAddresses: true selector: app: rec-infer ports: - name: peer port: 7351 --- apiVersion: v1 kind: Service metadata: name: rec-infer namespace: ranking spec: selector: app: rec-infer ports: - name: grpc port: 8471 targetPort: grpc ``` ## Checking it on a live cluster - List the peer Service's EndpointSlices with the `kubernetes.io/service-name` label and compare each endpoint's `ready` and `serving` conditions. With the flag working, an unready replica shows `ready: true` and `serving: false`. - Resolve a peer's per-pod name from inside another replica. - If a replica's endpoint is missing entirely, check that the pod has an IP and that the Service selector matches its labels. ## One more trap outside DNS Under the default StatefulSet `podManagementPolicy` of `OrderedReady`, the controller does not create `rec-infer-1` until `rec-infer-0` is Ready. If readiness needs a quorum, that alone is a deadlock that no DNS setting can fix. Quorum systems whose readiness depends on peers therefore usually run with `Parallel` pod management. How the StatefulSet orders pods belongs to StatefulSet design; the point here is to recognize both deadlocks, because they look alike from the outside. ## Deciding what readiness should mean The flag is a workaround for a design choice: readiness that depends on peers. Before reaching for it, ask what the readiness probe is for. If it means "this replica can answer ranking clients", a quorum check belongs in it, and the two-Service layout above is the right answer. If a replica can serve some traffic alone, a narrower probe removes the deadlock without any Service change. Either way, write down which Service each consumer uses, because the flag's effect is invisible from the pod side.
- What happens to a terminating replica's DNS record when publishNotReadyAddresses is set?It stays. The EndpointSlice controller keeps terminating pods in the slice, and with the flag set their `ready` condition is still true, so peers keep resolving a member that is shutting down. The `terminating` condition is set, but a DNS consumer that follows `ready` ignores it. The application's own membership protocol must detect the departure rather than rely on DNS.
- Why is the per-pod record still missing for a replica even with the flag set?An endpoint needs a pod IP, so a pod whose network sandbox is not set up yet is not listed at all. Terminal pods are dropped too. Beyond that, check that the Service selector matches the pod's labels and that the Service name equals the StatefulSet's `spec.serviceName`, since the hostname is only copied when the subdomain matches.
It is a club that only lists members in its directory once they have attended a meeting, while a meeting needs a quorum of listed members: the rule has to be relaxed for the founding members.
saying these in an interview costs you the question
- publishNotReadyAddresses makes the pods themselves pass their readiness probes
- Setting the flag on the client-facing Service is harmless
- With the flag set, a pod is published before it has an IP
- The flag makes cluster DNS drop terminating pods faster
- Fixing DNS always fixes the bootstrap, whatever podManagementPolicy is set