How does Kubernetes implement service discovery under the hood - what roles do the Service object, CoreDNS, and kube-proxy each play, and how does this differ from a registry-based tool like Consul or Eureka?
answer
- Service = stable ClusterIP + DNS name over a Pod selector
- EndpointSlice tracks ready Pod IPs
- readiness probe controls membership, not liveness probe
- CoreDNS resolves name; kube-proxy does the actual LB via iptables/IPVS
- no app-level registry client needed = server-side discovery
basics
~20 sKubernetes gives every Service a stable name and virtual IP. CoreDNS resolves the name to that IP, and kube-proxy sets up networking rules on each node so traffic to that IP gets sent to one of the healthy Pods behind it - no app-level registry client needed.
solid answer
~60 sA Kubernetes Service is a stable abstraction over a dynamic set of Pods selected by a label selector; it gets a persistent ClusterIP and a DNS name (service.namespace.svc.cluster.local) that CoreDNS resolves. The actual set of healthy Pod IPs backing a Service is tracked in Endpoints/EndpointSlice objects, which the Endpoints controller updates continuously based on each Pod's readiness probe status - a Pod failing its readiness probe is pulled out of EndpointSlices without being killed. kube-proxy, running as a daemon on every node, watches EndpointSlices and programs the node's iptables or IPVS rules so that any traffic sent to the Service's ClusterIP gets DNAT'ed and load-balanced to one of the current healthy Pod IPs, all in the kernel data path, with no per-request registry lookup happening in the app. The key contrast with Consul/Eureka is architectural: Kubernetes doesn't expose a separate registry API the app must query - discovery is baked into cluster networking itself via DNS + kube-proxy, making it a canonical server-side discovery implementation, whereas Consul/Eureka are typically used to support client-side discovery outside or across Kubernetes clusters, or for non-K8s and legacy workloads.
go deeper
Knows that Kubernetes services can be called by name and that this 'just works' without hardcoding IPs.
Can describe Service + selector + ClusterIP + DNS name at a high level.
Explains the full chain - readiness-driven EndpointSlice membership, CoreDNS resolution, kube-proxy's iptables/IPVS programming - and contrasts it with registry-based tools.
Reasons about convergence/consistency windows at scale (iptables churn, EndpointSlice sharding) and when to layer a mesh or use Consul for cross-cluster discovery.
## How Kubernetes differs from a standalone registry Kubernetes solves service discovery in a fundamentally different way than a standalone registry like Consul or Eureka: instead of exposing a registry API that application code queries at runtime, it bakes discovery directly into cluster networking, so that from the calling Pod's point of view, calling another service looks exactly like a normal DNS lookup plus a normal TCP connection - there's no discovery-specific client code at all. ## The Service object and readiness-driven membership The starting point is the Service object. A Service doesn't point at specific Pods directly; instead it defines a **label selector** (e.g. `app: orders`), and Kubernetes' control plane continuously matches that selector against all currently running Pods to figure out which ones back the Service. Crucially, that matching also filters on **readiness**: - a Pod only counts as a backend once it passes its configured readiness probe (an HTTP, TCP, or exec check the kubelet runs repeatedly); - the moment a Pod starts failing readiness - overloaded, warming up a cache, a dependency down - it's pulled out of the Service's backend set without the Pod itself being killed or restarted, which is a deliberate mechanism for load-shedding during transient trouble. This live backend set is materialized as `Endpoints` (legacy, one object per Service) or, in modern clusters, `EndpointSlice` objects (sharded for scale), each just a list of ready Pod IPs and ports; the Endpoints controller keeps these in sync with Pod readiness changes in near real time. ## Two mechanisms turn that list into a callable address Two separate mechanisms then turn that backend list into something an app can actually call. 1. **First**, every Service gets a name (`service-name.namespace.svc.cluster.local`, plus a shorter form resolvable within the same namespace) and a stable virtual IP called the **ClusterIP**, allocated once and never reused for the Service's lifetime even as the Pods behind it churn constantly. CoreDNS, running as cluster-internal DNS, answers lookups for that name with the ClusterIP - this is the discovery step from the calling code's perspective: a plain DNS resolution, exactly like resolving any public hostname, with no Kubernetes-specific API call in the application. 2. **Second**, `kube-proxy` - a daemon running on every single node in the cluster - continuously watches the Endpoints/EndpointSlice objects and, on any change, reprograms that node's networking rules (classically iptables DNAT rules; on more modern/larger clusters, IPVS, which scales better with very large numbers of Services). The effect is that any packet a Pod on that node sends to a Service's ClusterIP gets rewritten in the kernel to the address of one of the currently-ready backend Pods, chosen essentially at random - entirely inside the node's networking stack, before the packet even leaves the machine in many cases. So the full path for one Pod calling a Service is: DNS lookup resolves the name to the stable ClusterIP (fast, cacheable, essentially never changes), and the actual routing/load-balancing to a live Pod happens per-packet via kernel NAT rules that are kept current by kube-proxy watching real-time readiness state. ## Contrast with Consul and Eureka This is worth contrasting directly with Consul or Eureka. Those tools are registries with an explicit API surface: something must call register/deregister (or be heartbeat-monitored) and something must call the lookup API to get instance data - which is exactly the client-side discovery pattern, or requires an intermediary (a Consul-aware load balancer, or a proxy configured via Consul Connect) to act as the server-side piece. Kubernetes gives you server-side discovery for free, baked into the platform: no app code touches a registry API, and the mechanics (readiness-driven membership, DNS resolution, kernel-level load balancing) are entirely infrastructure-owned. This is precisely why running Consul or Eureka alongside Kubernetes is fairly rare for pure in-cluster service-to-service calls - Kubernetes' built-in Service discovery already covers that need - but those tools remain relevant for discovery that spans outside Kubernetes' boundary: - legacy VMs; - multi-cluster/multi-cloud topologies; - non-Kubernetes workloads that need to participate in the same discovery domain as containerized services. ## The convergence window One operational nuance worth knowing: kube-proxy's iptables mode does per-node, eventually-consistent rule updates - under very high EndpointSlice churn (say, a large Deployment rolling out) there can be a brief window where a newly dead Pod's IP is still in some node's rules until the next sync, which is one reason production setups often layer a service mesh's more actively-managed sidecar proxies on top when they need tighter control over that convergence window.
- What's the difference between a Pod's readiness probe and liveness probe, and which one controls whether it receives traffic via a Service?The liveness probe controls whether the kubelet restarts the container at all - failing it means Kubernetes thinks the process is broken and kills/restarts it. The readiness probe controls whether the Pod is included in the Service's EndpointSlice and therefore receives traffic; a Pod can fail readiness (e.g. still warming up, or a downstream dependency is down) without being restarted, simply getting excluded from load balancing until it passes readiness again.
- Why might kube-proxy's iptables mode become a scaling concern in very large clusters, and what's the common mitigation?iptables rule evaluation is roughly linear in the number of rules, and with thousands of Services and endpoints, packet processing latency and rule-update time can degrade. The common mitigation is switching kube-proxy to IPVS mode, which uses hash-table-based lookups that scale better with large numbers of Services, or moving to eBPF-based dataplanes that bypass iptables/IPVS entirely.
It's like a company's single published support phone number (the Service DNS name/ClusterIP) that never changes, while behind the scenes a call-routing switchboard (kube-proxy) is constantly updated with which specific staff members (Pods) are currently at their desk and ready to take calls (passing readiness checks).
saying these in an interview costs you the question
- thinks Kubernetes Pods have a registry API app code must call
- confuses readiness probes with liveness probes
- doesn't know DNS resolves to a stable ClusterIP rather than a Pod IP directly
- unaware that kube-proxy (not CoreDNS) does the actual load-balancing to a specific Pod