A Kubernetes Service exists and its backing pods are Running, but requests to the Service's cluster IP hang, and the EndpointSlice for that Service lists no ready addresses. What causes an empty endpoint list, and how do you track it down?
answer
- Service selector matches POD labels (template), not Deployment labels
- Selector is AND — an extra key matches nothing
- Namespace-scoped: no cross-namespace selection
- READY 0/1 → address present but ready:false → excluded
- Named targetPort must exist in container ports
basics
~20 sA Service only routes to pods whose labels match its selector and that are Ready. Empty endpoints means either the selector matches nothing (label typo, wrong namespace) or the matching pods fail their readiness probe. Compare the Service selector with actual pod labels, then check readiness.
solid answer
~60 sA Service is a selector plus a port mapping; the EndpointSlice controller fills in the addresses. Empty means the controller found nothing to add, and there are only a few reasons: 1. **Selector/label mismatch.** The Service's `selector` must match the *pod* labels — that is `spec.template.metadata.labels` on a Deployment, not the Deployment's own labels or its `matchLabels` by coincidence. A single typo or an extra key silences the whole Service. 2. **Wrong namespace.** A Service selects only pods in its own namespace. 3. **Pods not Ready.** Matching pods whose readiness probe fails are recorded with `ready: false`, so kube-proxy will not send them traffic. This is the most common case in practice, and the fix is in the app or the probe, not the Service. 4. **Named targetPort not declared** by the container, so no port can be resolved. 5. **Service with no selector at all** — then endpoints are managed manually or it is an ExternalName. Diagnose with `kubectl describe svc`, `kubectl get pods -l <selector> --show-labels`, and `kubectl get endpointslices -l kubernetes.io/service-name=<svc>`.
code
bash · 4 lineskubectl describe svc web -n team-a
kubectl get svc web -n team-a -o jsonpath='{.spec.selector}{"\n"}'
kubectl get pods -n team-a -l app=web --show-labels
kubectl get endpointslices -n team-a -l kubernetes.io/service-name=web -o yamlgo deeper
Know that a Service routes to pods matched by its selector and know the two commands: describe svc and get pods -l <selector> --show-labels.
Explain the three label locations on a Deployment, AND semantics of selectors, namespace scoping, and how readiness controls endpoint inclusion.
Read EndpointSlice conditions directly, distinguish 'no matching pods' from 'matching but not ready', and treat a failing readiness probe as a signal to fix rather than to suppress.
Frame endpoint population as the contract between workload health and traffic routing, and discuss how conventions on labels, probe design and manifest generation prevent this failure class fleet-wide.
## How a Service acquires backends A `Service` object of type ClusterIP contains a stable virtual IP, one or more port definitions, and a label `selector`. It does not contain pod addresses. A controller in the control plane continuously lists pods in the Service's namespace, keeps the ones whose labels match the selector, and writes their IPs into **EndpointSlice** objects (the modern, shardable replacement for the older single `Endpoints` object; both may be present because Endpoints are mirrored for compatibility). On every node, kube-proxy watches those slices and programs iptables or IPVS rules that DNAT traffic aimed at the cluster IP to one of the ready backend addresses. So "no endpoints" means the middle step produced nothing, and everything downstream — DNS, kube-proxy, NetworkPolicy — is irrelevant until it does. The symptom is usually a hang or immediate failure with no backend ever contacted. ## Cause 1: the selector does not match any pod labels This is where most people go wrong, because a Deployment has labels in three places: ```yaml metadata: labels: {app: web, managed-by: platform} # labels ON the Deployment object spec: selector: matchLabels: {app: web} # how the Deployment finds ITS pods template: metadata: labels: {app: web, version: v2} # labels ON THE PODS <-- the ones that matter ``` A Service's selector is matched against the third set. Copying labels from the Deployment's `metadata` is a classic error, as is adding a key to the Service selector (`app: web, tier: backend`) that the pod template does not carry — selectors are pure AND, so an extra key that no pod has matches nothing. Check it mechanically rather than by eye: ``` kubectl get svc web -o jsonpath='{.spec.selector}' kubectl get pods -l app=web,tier=backend --show-labels ``` If the second command returns nothing, you have found the bug. ## Cause 2: namespace scope Services select only within their own namespace. A Service in `frontend` cannot select pods in `backend`; cross-namespace access is done by addressing the other namespace's Service (`api.backend.svc.cluster.local`), or with an ExternalName Service. Manifests copied between namespaces frequently trip this. ## Cause 3: pods match but are not Ready This is the case that fools people, because `kubectl get pods` shows `Running` — but look at the READY column: `0/1` means the container is up and the readiness probe is failing. The EndpointSlice controller still records the address, with `conditions.ready: false`, and kube-proxy excludes it. Effectively the Service has no usable backends. ``` kubectl get endpointslice -l kubernetes.io/service-name=web -o yaml | grep -A3 conditions kubectl describe pod <pod> | grep -A5 Readiness ``` Causes of failing readiness are the usual probe problems: wrong path or port, the app bound to `127.0.0.1` so the kubelet cannot reach it, an aggressive `timeoutSeconds` (default 1s), or a genuine dependency the app is waiting on. Note the design intent: readiness failing *is* the system protecting you from routing to a pod that cannot serve. The bug is in the pod or the probe, and "fixing" it by deleting the readiness probe converts a visible outage into 500s for users. ## Cause 4: port resolution A Service port may specify `targetPort` as a **name** rather than a number. That name must exist in `ports[].name` on the container. If it does not, the port cannot be resolved for that pod and it contributes no endpoint for that port. Numeric `targetPort` values do not have this problem — the endpoint is created regardless of whether anything is listening, which is why a wrong number gives you endpoints plus connection-refused rather than empty endpoints. ## Cause 5: the Service has no selector A selectorless Service is legitimate — it is how you point a cluster-internal name at an external database by writing EndpointSlices yourself, and how `ExternalName` Services (which are pure DNS CNAMEs and have no endpoints by design) work. If someone removed the selector, no controller will populate anything. ## A repeatable checklist 1. `kubectl describe svc <name>` — read Selector, Port/TargetPort, and the Endpoints line. 2. `kubectl get pods -n <ns> -l <selector> --show-labels` — do any pods match at all? 3. If pods match, check the READY column and each pod's readiness probe. 4. `kubectl get endpointslices -l kubernetes.io/service-name=<name> -o yaml` — addresses present but `ready: false` confirms readiness as the cause. 5. Only once endpoints are ready and non-empty is it worth looking at DNS, ports or network policy. The ordering matters: it converts a vague "the service is down" into one of a handful of concrete states within a minute.
- The pods are Running and their labels match the Service selector, but the EndpointSlice lists the addresses with ready: false. What now?That means the readiness probe is failing, so kube-proxy deliberately excludes those pods. Check the READY column in kubectl get pods and the Readiness section of kubectl describe pod for the probe's path, port and error. Typical causes are a wrong path or port, an application bound to 127.0.0.1 so the kubelet cannot reach it, a one-second default timeout against a slower endpoint, or the app genuinely waiting on a dependency. Fix the probe or the app — removing the probe just hides the failure from Kubernetes and exposes it to users.
- What is the difference between the Endpoints object and EndpointSlice, and which should you inspect?Endpoints is the original single object per Service holding every backend address, which scales badly for large Services because any change rewrites the whole object and is pushed to every node. EndpointSlice splits the same information into multiple smaller objects and adds per-address conditions such as ready, serving and terminating, plus topology hints. Modern clusters use EndpointSlice as the source of truth and mirror a legacy Endpoints object for compatibility, so inspect EndpointSlices — they carry the readiness detail you need.
- Can a Kubernetes Service select pods in another namespace?No. The selector is evaluated only against pods in the Service's own namespace. To reach workloads elsewhere you address that namespace's own Service by its qualified DNS name, such as api.backend.svc.cluster.local, or create an ExternalName Service that resolves to it. A selectorless Service with hand-written EndpointSlices can point at arbitrary IPs, including pods elsewhere, but that bypasses the normal label-driven model and has to be maintained by hand.
The Service is a phone extension, the EndpointSlice is the list of desks it rings. If nobody's badge matches the extension's rule, or everyone at those desks has flipped their status to 'not available', the extension rings into silence — the phone system is fine.
saying these in an interview costs you the question
- Matching the Service selector against the Deployment's metadata labels instead of the pod template labels
- Assuming a Running pod is automatically a Service backend, ignoring the READY column
- Deleting or weakening the readiness probe to make endpoints appear
- Believing a Service can select pods in a different namespace
- Jumping straight to DNS or NetworkPolicy before confirming the Service has any ready endpoints