Why does a gRPC client calling a Kubernetes ClusterIP Service pin its calls to one pod, and how does a headless Service fix that?
answer
- connection versus request balancing
- HTTP/2 multiplexes on one connection
- clusterIP None, no kube-proxy rules
- resolver target plus round robin
- max connection age forces re-resolve
basics
~20 skube-proxy picks a backend once per connection, and gRPC sends every call over one long-lived HTTP/2 connection, so one pod gets all of them. A headless Service returns every ready pod's IP, letting the client balance calls itself.
solid answer
~40 sA ClusterIP Service is a virtual IP that kube-proxy's node rules rewrite to one pod **when a connection opens**; connection tracking keeps the connection there. gRPC multiplexes all calls over one long-lived HTTP/2 connection, so every call goes to that one pod. With `clusterIP: None` there are no kube-proxy rules and DNS returns one A record per ready pod. The client must then dial a resolver target like `dns:///rec-infer-headless.ranking.svc.cluster.local:8471` and use a **round-robin** policy, because a pick-first default still uses one address. The catch is staleness: clients typically re-resolve only when a connection breaks, so new pods stay idle. A server-side maximum connection age forces periodic reconnects and re-resolution.
code
bash · 1 linekubectl -n ranking run dnsprobe --rm -it --restart=Never --image=busybox:1.36 -- nslookup rec-infer-headless.ranking.svc.cluster.localgo deeper
Remember that a normal Service balances connections, gRPC reuses one connection, and a headless Service returns every pod IP instead of one virtual IP.
Explain the kube-proxy rewrite at connection setup, conntrack pinning, HTTP/2 multiplexing, and the client-side pieces: a resolver target and a round-robin policy.
Show you have debugged it: idle pods after scale-out, stale records, connections that outlive pod readiness, and a server-side maximum connection age as the fix.
Frame the choice between client-side balancing, the VIP and a proxy or mesh by who owns the client libraries, how many languages you support, and the cost of another hop.
## Why a virtual IP balances connections, not calls A normal **ClusterIP Service** gives clients one virtual IP (VIP). No process listens on it. On every node, **kube-proxy** programs packet rules (iptables, nftables or IPVS) that rewrite traffic sent to the VIP towards one backend pod. The choice is made **when a connection opens**. Connection tracking then keeps every later packet of that connection on the same pod. That is fine for short HTTP/1.1 requests. It is not fine for **gRPC**: - gRPC runs over **HTTP/2**, which carries many concurrent calls (streams) on one TCP connection. - A gRPC client usually opens that connection once and keeps it for hours. - kube-proxy made its choice once, for that one connection, so **every call lands on the same pod**. Picture a recommendation-model inference server, `rec-infer`, with 3 replicas on port 8471. Two ranking clients each open one connection. With 3 pods and 2 connections, at least one pod gets no traffic at all, and the pods that do get traffic can be very unevenly loaded. ## What a headless Service hands the client instead A **headless Service** sets `clusterIP: None`. kube-proxy skips such Services entirely, so there are no VIP rules. Cluster DNS answers the Service name with **one A (or AAAA) record per ready pod**. The client now sees the real pod addresses and can balance across them itself. ```bash kubectl -n ranking run dnsprobe --rm -it --restart=Never --image=busybox:1.36 -- nslookup rec-infer-headless.ranking.svc.cluster.local ``` With 3 ready replicas that lookup returns three pod IPs; the same lookup against the ClusterIP Service returns its single VIP. ## Making the client use the whole record set Switching the Service alone changes nothing if the client still connects to one address. You also need to: 1. **Dial a resolver target** such as `dns:///rec-infer-headless.ranking.svc.cluster.local:8471`, so the client resolves every address rather than one. 2. **Choose a balancing policy** that uses them all. Many gRPC implementations default to *pick first*, which connects to a single address; a *round robin* policy opens a subchannel to each pod and spreads **each call**. 3. **Make re-resolution happen.** gRPC clients typically re-resolve when a connection breaks or the server asks them to go away, not on a timer tied to the DNS TTL. A server-side **maximum connection age** forces clients to reconnect, and so re-resolve, every few minutes. 4. **Watch runtime DNS caches.** Some language runtimes (the JVM, for example) cache lookups on their own terms, on top of any resolver cache. ## Stale records: scale-out and restarts Headless DNS only reflects the cluster at the moment of the lookup. Two failure modes follow: - **Scale-out is invisible.** Scale `rec-infer` from 3 to 5 replicas, and long-lived clients that resolved before the scale keep using 3 pods. The 2 new pods can sit idle for as long as those connections live. - **Removal does not close connections.** When a pod stops being ready, it leaves the DNS answer, but a client already connected to it keeps its connection. Client-side health checking, or the server closing connections as it shuts down, handles that; DNS does not. | Concern | ClusterIP VIP | Headless + client balancing | |---|---|---| | Balancing unit | Connection | Call (with a round-robin policy) | | Who picks the pod | kube-proxy rules on the node | The client library | | New pods noticed | On the next new connection | On the next re-resolution | | Client complexity | None | Resolver, policy, reconnect tuning | | Works for any client | Yes | Only clients that can use several addresses | ## When to pick which - **Keep the VIP** for short-lived HTTP/1.1 traffic, or for clients you cannot configure. - **Use headless plus client-side balancing** for long-lived multiplexed connections (gRPC, some database drivers) where you control the client. - **Use an L7 proxy or a service mesh** when you want per-call balancing without touching every client; that moves the balancing into a separate component, with its own cost. One subtlety worth saying in an interview: a headless Service is still an ordinary Service object. Its selector, its EndpointSlices and pod readiness all still decide which addresses DNS returns. Only the virtual IP and kube-proxy are gone.
- How do you get newly added pods used without restarting the gRPC clients?Make reconnects routine. A server-side maximum connection age closes each connection gracefully after a set time, the client reconnects and re-resolves the headless name, and new pods join the balancing set. Keep any runtime DNS cache short, since the JVM and some others cache lookups on their own. Expect a lag equal to the connection age, not the DNS TTL.
- What does a DNS SRV query against the headless Service add?For a named port such as `grpc` over TCP, a query for `_grpc._tcp.rec-infer-headless.ranking.svc.cluster.local` returns the port number together with one target per ready pod. A client that understands SRV learns the port as well as the addresses, so the port is not hardcoded. Most gRPC setups ignore SRV and use A records plus a configured port.
- What does the ClusterIP route still give you that the headless route does not?Zero client work. Any client, including ones you cannot configure, gets a single stable address, and every new connection lands only on a ready endpoint. For short HTTP/1.1 requests that is already even enough. If you need per-call balancing without changing clients, an L7 proxy or a service mesh can provide it, at the cost of another component in the path.
saying these in an interview costs you the question
- kube-proxy balances each gRPC call across the Service's pods
- A headless Service balances load by rotating DNS answers for you
- When DNS drops a pod, existing connections to it close automatically
- gRPC clients re-resolve exactly when the DNS record's TTL expires
- Switching the Service to headless is enough without changing the client