How does the informer/watch-based discovery client keep its instance data fresh, and what are the scaling and staleness tradeoffs versus native Kubernetes Service DNS?
answer
- Informer = initial LIST + long WATCH -> local cache
- getInstances = in-memory cache read
- Eventually consistent -> still need retries/circuit breaker
- Per-pod watches don't scale -> Discovery Server consolidates
- Native DNS = L4/ClusterIP/no-RBAC vs client-side LB
basics
~20 sNewer versions use Kubernetes informers: a watch stream feeds a local in-memory cache of Services/Endpoints, so getInstances() is a fast cache read that updates as the cluster changes. It trades some staleness and apiserver watch load for client-side load balancing that plain Service DNS can't do.
solid answer
~50 sThe KubernetesInformerDiscoveryClient uses shared informers: it does an initial list of Services and Endpoints, then holds a watch that streams add/update/delete events into a local cache. getInstances() reads that cache — cheap and current within the watch's propagation delay. This means each discovery-enabled pod holds open watch connections to the apiserver, so at large fleet scale the aggregate watch load and memory matter; the Discovery Server pattern consolidates this into one watcher. Staleness is bounded but nonzero: a pod that just failed can briefly remain in the cache until the endpoint-removal event arrives, so callers still need retries/circuit breakers. Versus native Service DNS + kube-proxy (which is L4, single virtual IP, server-side balancing, and requires no library or RBAC), the informer client gives pod-level instance lists, per-instance metadata, and Spring Cloud LoadBalancer client-side algorithms — at the cost of RBAC, watches, and cache complexity.
code
java · 27 lines// Discovery gives pod-level instances; pair with resilience since the cache is eventually consistent.
@Configuration
class ResilientClientConfig {
@Bean
@LoadBalanced // uses KubernetesInformerDiscoveryClient's cached instance list
WebClient.Builder webClientBuilder() {
return WebClient.builder();
}
}
@Service
class InventoryClient {
private final WebClient.Builder lb;
InventoryClient(WebClient.Builder lb) { this.lb = lb; }
// A cached instance may be briefly stale -> retry/circuit-break around the call
@CircuitBreaker(name = "inventory", fallbackMethod = "fallback")
@Retry(name = "inventory")
Mono<String> stock(String sku) {
return lb.build().get()
.uri("http://inventory/stock/{sku}", sku)
.retrieve().bodyToMono(String.class);
}
Mono<String> fallback(String sku, Throwable t) { return Mono.just("UNKNOWN"); }
}go deeper
Not expected to go deep; awareness that the client caches data is enough.
Should know discovery reads are cached and eventually consistent, so retries matter.
Should explain the informer LIST+WATCH model and the DNS-vs-DiscoveryClient tradeoff.
Should architect for watch fan-out (Discovery Server, HA, namespace/label scoping), reason about staleness bounds and resilience, and justify DNS vs client-side LB per workload.
**Informers — what they are.** A Kubernetes **informer** is a client-side pattern (from the Kubernetes client libraries) that maintains a local, eventually-consistent replica of a set of API objects. It works in two phases: (1) an initial **LIST** to populate the cache, then (2) a long-lived **WATCH** stream that delivers incremental **ADD / UPDATE / DELETE** events, keeping the cache in sync without repeated polling. A **shared informer** lets multiple consumers share one watch/cache to avoid duplicate load. **How discovery uses it.** `KubernetesInformerDiscoveryClient` (the modern implementation, in both the official-client and fabric8 families) runs informers over **Services** and **Endpoints/EndpointSlices**. So `getServices()` and `getInstances(serviceId)` are **in-memory cache reads** — fast, no per-call API round-trip. When a pod scales up, fails readiness, or is deleted, the corresponding Endpoints UPDATE event flows through the watch and the cache reflects it within a short propagation window. **Freshness / staleness model.** The cache is **eventually consistent**, not instantaneous: - There is a delay between the real-world event (pod dies) and the Endpoints controller updating the object, plus the watch delivery delay. - During that window `getInstances()` may still return a dead pod, or briefly miss a just-ready one. - Consequence: **you must still handle call failures** — Spring Cloud LoadBalancer + retries, timeouts, and a circuit breaker (Resilience4j) are non-optional. Discovery freshness is a best-effort optimization, not a correctness guarantee. **Scaling considerations (the principal-level crux):** 1. **Watch fan-out.** If every app pod embeds a discovery client, each opens watch connections to the apiserver. With thousands of pods this is real load (connections, event fan-out, memory on both sides). 2. **Consolidation via Discovery Server.** The **Spring Cloud Kubernetes Discovery Server** runs the informers once and serves discovery over HTTP, collapsing N watchers into 1 and removing per-app RBAC. Tradeoff: it becomes a dependency/SPOF you must make HA. 3. **Namespace scoping.** `all-namespaces=true` widens each informer's watch to the whole cluster — more events, more memory, broader RBAC. Prefer namespace scope. 4. **Label filters** (`service-labels`) shrink the watched set and cache size. **When to prefer native Service DNS instead.** Plain Kubernetes networking gives you `http://service.namespace.svc.cluster.local` → the Service **ClusterIP**, with **kube-proxy** (iptables/IPVS) doing **L4, server-side** load balancing. This needs **no library, no RBAC, no watches** and is the simplest, most robust choice. You give up: pod-level instance enumeration, per-instance metadata, client-side algorithms (e.g., zone-aware, weighted), and the portable `DiscoveryClient` abstraction. **Rule of thumb:** if all you need is 'call service B and get balanced,' native DNS wins on simplicity; reach for Spring Cloud Kubernetes discovery when you specifically need client-side load balancing, instance metadata, subsetting, or code portability across Eureka/Kubernetes. **Additional gotchas at scale:** - **Thundering herd on apiserver restart:** many informers re-LIST simultaneously; the Discovery Server mitigates this. - **Cache cold-start:** the first calls after boot may see an incomplete list until the initial LIST completes — design startup ordering/readiness accordingly. - **EndpointSlice migration:** large Services are chunked into multiple EndpointSlices; ensure the client version and RBAC cover `discovery.k8s.io/endpointslices`. - **Metadata cardinality:** copying all labels/annotations into instance metadata can bloat the cache; prefix and prune. **Bottom line:** the informer model buys cheap, near-real-time client-side discovery, but it's an eventually-consistent cache with real apiserver cost — architect for staleness (resilience) and for watch fan-out (Discovery Server) as the fleet grows.
- You're running 3,000 pods, each with an embedded discovery client, and the apiserver is straining on watch connections. What do you change?Move to the Spring Cloud Kubernetes Discovery Server: run the informers once in a (HA) deployment that serves discovery over HTTP, so the 3,000 app pods query it instead of each opening apiserver watches. This collapses watch fan-out and removes per-app RBAC. Also scope to namespaces and apply label filters.
- Why can't you rely on the discovered instance list being perfectly accurate at the moment of a call?The informer cache is eventually consistent — there's a delay from real event to Endpoints update to watch delivery. A just-failed pod can linger briefly. So you design for it with timeouts, retries against another instance, and a circuit breaker, treating discovery freshness as best-effort.
saying these in an interview costs you the question
- Claiming the instance cache is strongly consistent / always accurate
- Assuming embedded per-pod watches scale infinitely with no apiserver cost
- Believing native Service DNS gives client-side per-pod load balancing (it's L4 via kube-proxy)
- Skipping retries/circuit breakers because 'discovery is exact'