skip to content

How does Spring Cloud Kubernetes turn a Service into a list of ServiceInstances, and how do readiness, ports, and metadata factor in?

level: seniorimportance: should knowfreq 40%

answer

  1. Service -> Endpoints -> subsets.addresses
  2. Ready addresses only (readiness gates it)
  3. include-not-ready-addresses to reverse
  4. primary-port-name for multi-port Services
  5. metadata from labels/annotations/ports

basics

~20 s

For a serviceId it finds the matching Kubernetes Service and reads its Endpoints. Each ready pod address in the Endpoints subsets becomes one ServiceInstance, with pod IP, chosen port, and metadata built from the Service's labels, annotations, and ports.

solid answer

~40 s

getInstances(serviceId) resolves the Kubernetes Service by name, then reads its Endpoints object. Endpoints group addresses into subsets: 'addresses' (ready pods) and 'notReadyAddresses' (failing readiness). By default only ready addresses become ServiceInstances, so readiness probes directly control discoverability; you can opt into not-ready ones via include-not-ready-addresses. Each address yields a ServiceInstance whose host is the pod IP and whose port is chosen from the Service's ports — when several exist you disambiguate with primary-port-name or a port-name convention. Metadata is assembled from configurable sources: Service labels, Service annotations, and port names, each toggled and optionally prefixed via spring.cloud.kubernetes.discovery.metadata.* This ServiceInstance list feeds Spring Cloud LoadBalancer for client-side balancing across pods. Because Kubernetes updates Endpoints as pods scale or fail probes, the discovered list tracks the live topology without any heartbeat from the app.

code

java · 25 lines
java
@Component
class InstanceInspector {
    private final DiscoveryClient discoveryClient;
    InstanceInspector(DiscoveryClient discoveryClient) { this.discoveryClient = discoveryClient; }

    void dump() {
        for (ServiceInstance si : discoveryClient.getInstances("inventory")) {
            // host = pod IP, port = resolved (primary-port-name aware)
            System.out.printf("%s:%d meta=%s%n", si.getHost(), si.getPort(), si.getMetadata());
            // meta may include labels/annotations/port-names per discovery.metadata.* config
        }
    }
}

// application.yml
// spring:
//   cloud:
//     kubernetes:
//       discovery:
//         primary-port-name: http
//         include-not-ready-addresses: false
//         metadata:
//           labels:      { enabled: true, prefix: 'l_' }
//           annotations: { enabled: true, prefix: 'a_' }
//           ports:       { enabled: true }

go deeper

for a junior

Know that ready pods behind a Service become instances; details of ports/metadata optional.

for a middle

Should explain readiness gating and that pod IPs (not ClusterIP) are the instances.

for a senior

Should cover subsets, multi-port selection via primary-port-name, and metadata sources, plus LoadBalancer integration.

for a principal

Should reason about informer-cache freshness, headless services, cross-namespace RBAC, and failure modes of empty/misselected Endpoints.

**From name to instances — the mechanism.** When code (or Spring Cloud LoadBalancer) calls `getInstances("inventory")`: 1. Spring Cloud Kubernetes finds the **Service** named `inventory` (in the configured namespace, or across all namespaces if `all-namespaces=true`). 2. It reads the associated **Endpoints** object. In Kubernetes, an Endpoints object contains **subsets**, and each subset has: - `addresses` — pod IP:port pairs that are **ready** (passing their readiness probe), - `notReadyAddresses` — pods that exist but are not ready, - `ports` — the named ports for that subset. 3. Each entry in `addresses` becomes a `ServiceInstance`: `getHost()` = pod IP, `getPort()` = the resolved port, plus metadata. (Newer clusters use **EndpointSlice**; the concept is the same.) **Readiness gates discoverability.** This is the crucial behavioral point. A pod only lands in `addresses` once its **readiness probe** passes; a failing probe moves it to `notReadyAddresses`, and by default Spring Cloud Kubernetes **excludes** those. So a pod that is up but not ready (warming caches, failing a dependency check) is automatically *not* discovered — no traffic. `spring.cloud.kubernetes.discovery.include-not-ready-addresses=true` reverses this, which you almost never want except for diagnostics. **Port selection.** A Service can expose **multiple named ports** (e.g., `http`, `grpc`, `management`). Spring must pick one for the `ServiceInstance` port. Controls: - `spring.cloud.kubernetes.discovery.primary-port-name` — global default port name to prefer. - A `primary-port-name` annotation on the Service can override per-service. - If exactly one port exists, it's used; ambiguity without a hint can lead to picking an unexpected port — a common gotcha. **Metadata assembly.** Each `ServiceInstance.getMetadata()` is a `Map<String,String>` built from configurable sources under `spring.cloud.kubernetes.discovery.metadata`: - `labels.enabled` / `labels.prefix` — copy Service **labels** into metadata (optionally prefixed to avoid key clashes). - `annotations.enabled` / `annotations.prefix` — copy Service **annotations**. - `ports.enabled` / `ports.prefix` — expose the Service's port names→numbers. This metadata drives things like Spring Cloud LoadBalancer hints, zone/affinity routing, and custom instance filtering. **Integration with load balancing.** The returned `List<ServiceInstance>` is exactly what **Spring Cloud LoadBalancer** consumes. A `@LoadBalanced RestTemplate`/`WebClient` or Feign call to `http://inventory/...` triggers a `getInstances("inventory")`, LoadBalancer applies its algorithm (round-robin by default) across the *pods*, and the request goes straight to a pod IP — true client-side load balancing, unlike plain Service DNS which hits the ClusterIP and lets kube-proxy balance at L4. **Liveness of the list.** Because Kubernetes' Endpoints controller keeps Endpoints in sync with pod state and readiness, the discovered set reflects scaling and failures automatically. There is no app-side heartbeat. In informer-based implementations, a **watch** pushes changes into a local cache so `getInstances` is a fast in-memory read rather than a live API call each time. **Gotchas:** - A Service with a **selector that matches nothing** yields empty Endpoints → empty instance list. - **Headless Services** (`clusterIP: None`) still have Endpoints and work for pod-level discovery. - Not distinguishing named ports leads to calling the wrong port. - Cross-namespace discovery requires `all-namespaces` plus RBAC covering those namespaces.

  • A pod is Running but receives no traffic through discovery. What is the most likely cause?
    Its readiness probe is failing, so Kubernetes keeps it in notReadyAddresses rather than addresses; Spring Cloud Kubernetes excludes not-ready addresses by default, so it never appears in getInstances().
  • The client keeps hitting the management port instead of http. Why, and how do you fix it?
    The Service exposes multiple named ports and no primary was specified, so an unintended one was chosen. Set spring.cloud.kubernetes.discovery.primary-port-name (or the per-Service primary-port-name annotation) to 'http'.

saying these in an interview costs you the question

  • Saying instances come from the Service ClusterIP rather than Endpoints/pods
  • Claiming not-ready pods are discovered by default
  • Assuming there's always exactly one port so no selection is needed
  • Believing the app must poll/heartbeat to stay in the list

context