A Kubernetes Service of type ClusterIP gets an IP address that no network interface owns and that nothing answers ARP for. Explain what actually happens to a packet sent to that IP, and which component makes it work.
answer
- ClusterIP = VIP, no interface, no ARP
- kube-proxy programs rules, does not forward
- DNAT + conntrack reverse-translation
- Per-connection, not per-packet, L4 only
- No endpoints -> REJECT, not hang
basics
~20 sA ClusterIP is a virtual IP with no machine behind it. kube-proxy runs on every node and programs kernel rules that rewrite the destination of packets aimed at that IP to one of the Service's backing pod IPs (DNAT). A backend is chosen once per new connection; connection tracking rewrites the replies back.
solid answer
~50 sA ClusterIP is a **virtual IP**: allocated from the cluster's Service CIDR, never assigned to an interface, never answers ping or ARP. It exists only as a match rule in each node's kernel dataplane. **kube-proxy** runs on every node (normally as a DaemonSet), watches the API server for Services and EndpointSlices, and translates them into kernel rules - iptables chains, nftables rules, or IPVS virtual servers depending on mode. When a pod connects to `10.96.0.10:80`, the rule matches on egress, performs **DNAT** to a chosen backend pod IP and target port, and conntrack records the translation so replies are rewritten back to the ClusterIP. The client never sees the backend address. Selection happens **once per connection**, not per packet, and is effectively random or round-robin with no L7 awareness. With no ready endpoints the packet is rejected rather than DNATed. Note kube-proxy is a control agent for the dataplane - it is not itself on the packet path.
code
bash · 6 lineskubectl get svc web -o jsonpath='{.spec.clusterIP}{"\n"}'
kubectl get endpointslices -l kubernetes.io/service-name=web
# on a node (iptables mode)
sudo iptables -t nat -L KUBE-SERVICES -n | grep 10.96.0.10
sudo conntrack -L -d 10.96.0.10go deeper
Be able to say: ClusterIP is a virtual IP, kube-proxy on each node writes kernel rules, packets are DNATed to a pod IP, and the backend is picked per connection.
Add the mechanics: which kernel subsystem does the rewrite, the role of conntrack for reply traffic and flow stickiness, and what happens when there are zero ready endpoints.
Draw the consequences for real traffic - long-lived HTTP/2 connections pinning to one pod, stale rules when kube-proxy stops syncing, and when you reach for an L7 proxy instead.
Frame it as an abstraction boundary: Service semantics are the contract, kube-proxy is one implementation, and eBPF kube-proxy replacements or mesh sidecars swap the implementation without changing workload code.
## The ClusterIP is deliberate fiction When you create a Service, the API server allocates an address from the **Service CIDR**, a range configured on the control plane (commonly something like `10.96.0.0/12`). That allocation is a record in etcd and nothing more. No node configures the address on an interface, no ARP or NDP responder claims it, and pinging it usually fails even when the Service works perfectly. The address is a *label the kernel matches on*, not a destination that exists. ## kube-proxy's job kube-proxy runs on every node. It watches the API server for two kinds of object: - **Services**, which give it the virtual IP, ports, protocol and session-affinity settings. - **EndpointSlices**, which give it the current set of backend pod IPs and their readiness. From those it computes the rules the node's kernel needs and writes them into the chosen dataplane. It then reconciles: every time endpoints change, it recomputes and re-syncs. Crucially, kube-proxy does not forward traffic. Once the rules are installed the kernel does all the work at line rate, which is why a crashed kube-proxy does not immediately break existing Services - it just stops them from being updated, so the rules go stale. ## What happens to the packet Say a pod at `10.244.1.7` connects to a Service at `10.96.0.10:80` whose backends are `10.244.2.3:8080` and `10.244.3.9:8080`. 1. The pod's socket sends a SYN to `10.96.0.10:80`. Routing sends it toward the node's stack. 2. In the NAT path (iptables `nat` OUTPUT/PREROUTING, or the IPVS/nftables equivalent) a rule matches destination `10.96.0.10` and TCP port 80. 3. One backend is selected. In iptables mode this is done with a chain of `statistic mode random probability` jumps that produce an approximately even split; in IPVS mode a real scheduler (round robin by default) picks it. 4. **DNAT** rewrites destination address and port to `10.244.2.3:8080`. The source is untouched for in-cluster traffic, so the backend sees the real client pod IP. 5. The kernel's **conntrack** table stores the tuple with its translation. Every subsequent packet of that flow follows the same entry - hence per-connection, not per-packet, load balancing - and reply packets get reverse-NATed so the client sees replies from `10.96.0.10:80`, which is what the client's socket expects. ## Consequences that show up in interviews - **Load balancing is L4 and per-connection.** A client that opens one long-lived HTTP/2 or gRPC connection pins to one backend forever. Fixing that needs client-side balancing, a headless Service, or a service mesh / L7 proxy - not kube-proxy. - **`ping` a ClusterIP fails** because rules typically match only the declared protocol and ports. Testing with `curl <clusterip>:<port>` is the correct check. - **No ready endpoints** means the rule set has nothing to DNAT to; iptables mode installs a REJECT so you get connection refused fast rather than a hang. - **Session affinity** (`service.spec.sessionAffinity: ClientIP`) is the only stickiness kube-proxy offers, implemented with a client-IP timeout in the same kernel machinery. - **Headless Services** (`clusterIP: None`) opt out of all of this: no virtual IP, no DNAT, DNS just returns pod IPs directly. ## Where this fits The pod-to-pod substrate (routing pod IPs between nodes) is the CNI plugin's job. kube-proxy sits on top: it turns a stable Service name and virtual IP into one of those pod IPs. Some modern CNIs (Cilium in kube-proxy-replacement mode, for example) implement the same Service semantics in eBPF and let you delete kube-proxy entirely - the abstraction stays identical from the workload's point of view.
- If kube-proxy crashes on a node, do existing Service connections on that node break?No. The rules are already in the kernel and keep forwarding traffic, and established flows stay pinned through conntrack. What breaks is convergence: endpoint changes stop being applied, so the node keeps sending traffic to pods that have been deleted and never learns about new ones. It degrades into stale routing rather than an outage, which makes it easy to miss without monitoring on kube-proxy's sync latency.
- Why does a single long-lived gRPC client end up hammering one pod even though the Service has ten?Because kube-proxy balances per TCP connection, not per request. gRPC multiplexes many requests over one HTTP/2 connection, so once that connection is DNATed to a backend, every request rides the same conntrack entry. Fixes are client-side load balancing over a headless Service, periodic connection recycling with a max-connection-age, or an L7 proxy / service mesh that balances per request.
The ClusterIP is like a phone extension that rings no physical handset: the switchboard rules rewrite each incoming call to whichever real desk phone is free, and the caller never learns the desk number.
saying these in an interview costs you the question
- Claiming kube-proxy proxies the traffic itself in userspace (the userspace mode was removed long ago; modern modes are pure kernel rules)
- Saying the ClusterIP is assigned to a node or to a load balancer somewhere
- Believing load balancing is per request, so gRPC/HTTP2 spreads automatically
- Concluding a Service is broken because ping to the ClusterIP fails
- Thinking DNS resolution to the ClusterIP is the load-balancing step (CoreDNS returns one stable VIP; balancing happens in the kernel)