skip to content

kube-proxy and EndpointSlices

A Service's ClusterIP is virtual - kube-proxy programs iptables or IPVS rules on every node to rewrite traffic to the pod IPs tracked in EndpointSlices. Explaining that rewrite is the classic 'how does a Service actually work?' answer.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

5

A Kubernetes Service of type ClusterIP gets an IP address that no network interface owns and that nothing answers ARP for. Explain what actually happens to a packet sent to that IP, and which component makes it work.

level: juniorimportance: must knowfreq 78%

answer

  1. ClusterIP = VIP, no interface, no ARP
  2. kube-proxy programs rules, does not forward
  3. DNAT + conntrack reverse-translation
  4. Per-connection, not per-packet, L4 only
  5. No endpoints -> REJECT, not hang

basics

~20 s

A ClusterIP is a virtual IP with no machine behind it. kube-proxy runs on every node and programs kernel rules that rewrite the destination of packets aimed at that IP to one of the Service's backing pod IPs (DNAT). A backend is chosen once per new connection; connection tracking rewrites the replies back.

solid answer

~50 s

A ClusterIP is a **virtual IP**: allocated from the cluster's Service CIDR, never assigned to an interface, never answers ping or ARP. It exists only as a match rule in each node's kernel dataplane. **kube-proxy** runs on every node (normally as a DaemonSet), watches the API server for Services and EndpointSlices, and translates them into kernel rules - iptables chains, nftables rules, or IPVS virtual servers depending on mode. When a pod connects to `10.96.0.10:80`, the rule matches on egress, performs **DNAT** to a chosen backend pod IP and target port, and conntrack records the translation so replies are rewritten back to the ClusterIP. The client never sees the backend address. Selection happens **once per connection**, not per packet, and is effectively random or round-robin with no L7 awareness. With no ready endpoints the packet is rejected rather than DNATed. Note kube-proxy is a control agent for the dataplane - it is not itself on the packet path.

code

bash · 6 lines
bash
kubectl get svc web -o jsonpath='{.spec.clusterIP}{"\n"}'
kubectl get endpointslices -l kubernetes.io/service-name=web

# on a node (iptables mode)
sudo iptables -t nat -L KUBE-SERVICES -n | grep 10.96.0.10
sudo conntrack -L -d 10.96.0.10

go deeper

for a junior

Be able to say: ClusterIP is a virtual IP, kube-proxy on each node writes kernel rules, packets are DNATed to a pod IP, and the backend is picked per connection.

for a middle

Add the mechanics: which kernel subsystem does the rewrite, the role of conntrack for reply traffic and flow stickiness, and what happens when there are zero ready endpoints.

for a senior

Draw the consequences for real traffic - long-lived HTTP/2 connections pinning to one pod, stale rules when kube-proxy stops syncing, and when you reach for an L7 proxy instead.

for a principal

Frame it as an abstraction boundary: Service semantics are the contract, kube-proxy is one implementation, and eBPF kube-proxy replacements or mesh sidecars swap the implementation without changing workload code.

## The ClusterIP is deliberate fiction When you create a Service, the API server allocates an address from the **Service CIDR**, a range configured on the control plane (commonly something like `10.96.0.0/12`). That allocation is a record in etcd and nothing more. No node configures the address on an interface, no ARP or NDP responder claims it, and pinging it usually fails even when the Service works perfectly. The address is a *label the kernel matches on*, not a destination that exists. ## kube-proxy's job kube-proxy runs on every node. It watches the API server for two kinds of object: - **Services**, which give it the virtual IP, ports, protocol and session-affinity settings. - **EndpointSlices**, which give it the current set of backend pod IPs and their readiness. From those it computes the rules the node's kernel needs and writes them into the chosen dataplane. It then reconciles: every time endpoints change, it recomputes and re-syncs. Crucially, kube-proxy does not forward traffic. Once the rules are installed the kernel does all the work at line rate, which is why a crashed kube-proxy does not immediately break existing Services - it just stops them from being updated, so the rules go stale. ## What happens to the packet Say a pod at `10.244.1.7` connects to a Service at `10.96.0.10:80` whose backends are `10.244.2.3:8080` and `10.244.3.9:8080`. 1. The pod's socket sends a SYN to `10.96.0.10:80`. Routing sends it toward the node's stack. 2. In the NAT path (iptables `nat` OUTPUT/PREROUTING, or the IPVS/nftables equivalent) a rule matches destination `10.96.0.10` and TCP port 80. 3. One backend is selected. In iptables mode this is done with a chain of `statistic mode random probability` jumps that produce an approximately even split; in IPVS mode a real scheduler (round robin by default) picks it. 4. **DNAT** rewrites destination address and port to `10.244.2.3:8080`. The source is untouched for in-cluster traffic, so the backend sees the real client pod IP. 5. The kernel's **conntrack** table stores the tuple with its translation. Every subsequent packet of that flow follows the same entry - hence per-connection, not per-packet, load balancing - and reply packets get reverse-NATed so the client sees replies from `10.96.0.10:80`, which is what the client's socket expects. ## Consequences that show up in interviews - **Load balancing is L4 and per-connection.** A client that opens one long-lived HTTP/2 or gRPC connection pins to one backend forever. Fixing that needs client-side balancing, a headless Service, or a service mesh / L7 proxy - not kube-proxy. - **`ping` a ClusterIP fails** because rules typically match only the declared protocol and ports. Testing with `curl <clusterip>:<port>` is the correct check. - **No ready endpoints** means the rule set has nothing to DNAT to; iptables mode installs a REJECT so you get connection refused fast rather than a hang. - **Session affinity** (`service.spec.sessionAffinity: ClientIP`) is the only stickiness kube-proxy offers, implemented with a client-IP timeout in the same kernel machinery. - **Headless Services** (`clusterIP: None`) opt out of all of this: no virtual IP, no DNAT, DNS just returns pod IPs directly. ## Where this fits The pod-to-pod substrate (routing pod IPs between nodes) is the CNI plugin's job. kube-proxy sits on top: it turns a stable Service name and virtual IP into one of those pod IPs. Some modern CNIs (Cilium in kube-proxy-replacement mode, for example) implement the same Service semantics in eBPF and let you delete kube-proxy entirely - the abstraction stays identical from the workload's point of view.

  • If kube-proxy crashes on a node, do existing Service connections on that node break?
    No. The rules are already in the kernel and keep forwarding traffic, and established flows stay pinned through conntrack. What breaks is convergence: endpoint changes stop being applied, so the node keeps sending traffic to pods that have been deleted and never learns about new ones. It degrades into stale routing rather than an outage, which makes it easy to miss without monitoring on kube-proxy's sync latency.
  • Why does a single long-lived gRPC client end up hammering one pod even though the Service has ten?
    Because kube-proxy balances per TCP connection, not per request. gRPC multiplexes many requests over one HTTP/2 connection, so once that connection is DNATed to a backend, every request rides the same conntrack entry. Fixes are client-side load balancing over a headless Service, periodic connection recycling with a max-connection-age, or an L7 proxy / service mesh that balances per request.

The ClusterIP is like a phone extension that rings no physical handset: the switchboard rules rewrite each incoming call to whichever real desk phone is free, and the caller never learns the desk number.

saying these in an interview costs you the question

  • Claiming kube-proxy proxies the traffic itself in userspace (the userspace mode was removed long ago; modern modes are pure kernel rules)
  • Saying the ClusterIP is assigned to a node or to a load balancer somewhere
  • Believing load balancing is per request, so gRPC/HTTP2 spreads automatically
  • Concluding a Service is broken because ping to the ClusterIP fails
  • Thinking DNS resolution to the ClusterIP is the load-balancing step (CoreDNS returns one stable VIP; balancing happens in the kernel)

context

open as a page

Kubernetes' kube-proxy component can run in iptables mode or IPVS mode. Compare how each implements Service load balancing, and explain what actually changes as the number of Services and endpoints grows.

level: middleimportance: must knowfreq 62%

basics

~20 s

iptables mode builds a linear chain of NAT rules per Service and picks a backend with random-probability jumps; rule count grows with Services times endpoints, and every change means re-syncing a large ruleset. IPVS mode uses the kernel's L4 load balancer with hash-table lookups, so matching is near constant time, updates are incremental, and real schedulers (round-robin, least-connection, source-hash) are available.

open as a page

During a rolling update of a Kubernetes Deployment, clients see a burst of connection-refused and reset errors even though the app implements graceful shutdown on SIGTERM. Walk through why this happens and how you would eliminate it.

level: seniorimportance: must knowfreq 58%

basics

~20 s

Pod deletion and endpoint removal are concurrent, not ordered. The kubelet can send SIGTERM and the app can stop listening before every node's kube-proxy has removed that pod from its rules, so in-flight traffic still arrives. Fix it by making the container keep serving briefly after deletion starts - a preStop sleep or a delayed shutdown - so proxies converge before the socket closes.

open as a page

Kubernetes replaced the older Endpoints API with EndpointSlices for tracking the backends of a Service. What problem did that solve, and how does the newer object differ in structure?

level: middleimportance: should knowfreq 48%

basics

~20 s

The old Endpoints object crammed every backend of a Service into one resource, so one pod change rewrote the whole object and pushed it to every watcher - painful at thousands of endpoints. EndpointSlices shard the same data into capped chunks (100 endpoints by default), so a change updates one small slice, and they add per-endpoint conditions and topology fields.

open as a page

You run a multi-tenant Kubernetes platform heading toward tens of thousands of Services and very high pod churn. How would you decide between keeping kube-proxy in iptables mode, moving to IPVS or nftables, or removing kube-proxy in favour of an eBPF dataplane?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Decide on measured convergence and failure behaviour, not benchmarks. Keep iptables while endpoint-sync latency stays inside your rollout error budget; move to nftables or IPVS when sync time is the bottleneck; drop kube-proxy for eBPF only when you accept coupling the Service dataplane to one CNI and can staff kernel-level debugging.

open as a page