skip to content

What is kube-proxy's job on a Kubernetes node, and what still works if kube-proxy stops running there?

level: middleimportance: should knowfreq 50%

answer

  1. watches Services + EndpointSlices, writes kernel rules
  2. iptables / IPVS / nftables; userspace mode is dead
  3. controller not data-path proxy; kernel does the DNAT
  4. dies -> stale rules persist, no convergence
  5. not CNI, not DNS, not Ingress; L4 only

basics

~20 s

kube-proxy runs on every node (usually as a DaemonSet) and programs the node's kernel so traffic to a Service's virtual IP is redirected to a healthy backend Pod. It is a control agent writing rules, not a data-path proxy: if it stops, existing rules keep working but new Services and endpoint changes stop being applied.

solid answer

~50 s

kube-proxy is the node-level agent that makes **Service ClusterIPs work**. It watches Services and EndpointSlices from the API server and translates them into kernel rules on that node — iptables chains by default, or IPVS virtual servers in `ipvs` mode. When a Pod connects to a ClusterIP, the kernel DNATs the packet to one of the ready backend Pod IPs; there is no per-packet userspace hop (the old `userspace` mode is long gone). It also programs NodePort listeners and implements `externalTrafficPolicy`, source-NAT decisions and session affinity. If kube-proxy dies, **already-programmed rules stay in the kernel**, so existing Services keep resolving to the backends they had. What breaks is convergence: new Services get no rules, rescheduled or scaled Pods are not picked up, and endpoints of dead Pods are never removed, so some traffic hits black holes. Note that some CNI dataplanes (Cilium's kube-proxy replacement, for example) implement Service handling in eBPF and remove kube-proxy entirely.

code

bash · 11 lines
bash
kubectl -n kube-system get ds kube-proxy -o wide
kubectl -n kube-system logs -l k8s-app=kube-proxy --tail=50

# iptables mode
iptables-save -t nat | grep 10.96.0.42

# ipvs mode
ipvsadm -Ln

# compare with what the control plane says the backends are
kubectl get endpointslices -l kubernetes.io/service-name=my-svc

go deeper

for a junior

Say that kube-proxy runs on each node and makes Service ClusterIPs reach the right Pods by programming rules in the node's kernel.

for a middle

Explain that it watches Services and EndpointSlices and writes iptables or IPVS rules, and that the kernel — not kube-proxy — forwards the packets.

for a senior

Reason about failure modes: stale rules after kube-proxy dies, partial timeouts, externalTrafficPolicy and source-NAT effects, and how to compare programmed rules against EndpointSlices.

for a principal

Weigh dataplane choices — iptables versus IPVS versus nftables versus an eBPF kube-proxy replacement — against Service count, latency targets, observability and operational coupling to the CNI.

## What problem it solves A Service gets a stable virtual IP (the ClusterIP) that no network interface actually owns. Something has to make packets sent to that IP land on a real Pod. That something is the node's kernel, and kube-proxy is the agent that keeps the kernel's rules in sync with the API. The key mental model: **kube-proxy is a controller, not a proxy.** Despite the name, in the default modes it does not sit in the data path. It watches, it writes rules, and the kernel forwards at line rate. ## What it watches and writes kube-proxy watches two resources: **Services** (ClusterIP, ports, session affinity, externalTrafficPolicy) and **EndpointSlices** (the ready backend Pod IPs, maintained by the EndpointSlice controller based on readiness). On any change it recomputes and applies the node's rules. Modes: - **iptables** (default): chains that DNAT the ClusterIP and port to a randomly chosen backend, with connection tracking pinning a flow to its backend. Simple and robust; rule-set size grows with Services times endpoints, and full resyncs get expensive in very large clusters. - **IPVS**: uses the kernel's L4 load balancer with hash-table lookups, scaling better to thousands of Services and offering scheduling algorithms (round-robin, least-connection, hashing). Still uses some iptables rules for masquerading and filtering. - **nftables**: a newer backend replacing the iptables one on modern clusters, with better scaling characteristics. - **userspace**: the original mode, long deprecated and removed — describing it as current is a red flag. It also handles: - **NodePort**: making the allocated port on every node forward into the Service. - **externalTrafficPolicy**: `Cluster` load-balances to any backend but source-NATs (losing the client IP); `Local` forwards only to backends on the same node, preserving client IP and failing health checks on nodes with no backend. - **Session affinity** (`sessionAffinity: ClientIP`) and masquerading for traffic that must appear to come from the node. ## What kube-proxy does NOT do Being precise here separates candidates: - It does **not** assign Pod IPs or set up pod networking — that is the CNI plugin, invoked at sandbox creation. - It does **not** resolve DNS names; a Service's DNS name is CoreDNS's job. kube-proxy only handles the IP the DNS answer contains. - It does **not** run Ingress or do HTTP-aware routing — it is L4 only: no path-based routing, no retries, no TLS termination. - It does **not** decide readiness. Endpoints appear and disappear because the kubelet's probes drive Pod readiness and the EndpointSlice controller reflects it; kube-proxy just consumes the result. ## Failure behaviour Because the rules live in the kernel, kube-proxy going down is a **degradation, not an outage**: - Existing Services keep working against the endpoint set that was current when it died. - New Services and NodePorts get no rules on that node — connections fail. - Endpoint changes are not applied, so Pods that moved elsewhere are still targeted and a share of connections is DNATed to IPs that no longer serve. That produces the classic partial, intermittent symptom: "about a third of requests time out, but only from one node". - Conversely, newly created Pods never receive traffic even though they are Ready. Debugging means checking that the kube-proxy Pod is running on that node, reading its logs, and inspecting the rules with `iptables-save` or `ipvsadm -Ln`. Comparing `kubectl get endpointslices` against what the node has programmed settles whether the problem is control-plane (endpoints wrong) or node-level (rules stale). ## The eBPF alternative Modern CNIs increasingly replace kube-proxy: Cilium's kube-proxy replacement implements ClusterIP, NodePort and load-balancer handling in eBPF programs attached at the socket and driver level, skipping iptables and conntrack entirely. That yields lower latency and better scaling, at the cost of tighter coupling to the CNI and stricter kernel requirements. Knowing kube-proxy is replaceable — and why anyone would — is a good senior signal.

  • A Service intermittently times out from one node only, while other nodes are fine. How does kube-proxy fit into your diagnosis?
    Per-node symptoms point at that node's programmed rules being stale, so first check whether the kube-proxy Pod is running and healthy there and read its logs. Then compare the node's iptables or IPVS entries for that ClusterIP against the current EndpointSlice — if the node still targets Pod IPs that no longer exist, kube-proxy has stopped converging. If the rules match the endpoints, look elsewhere: CNI, node routing, or the backend itself.
  • Why can a cluster run with no kube-proxy at all?
    Because kube-proxy is just one implementation of Service semantics. CNIs such as Cilium can implement ClusterIP, NodePort and load-balancer handling in eBPF attached at the socket and driver level, bypassing iptables and conntrack. The Service API is unchanged; only the dataplane differs, usually with lower latency and better scaling in large clusters, at the cost of tighter CNI coupling and kernel version requirements.

kube-proxy is the clerk who keeps the building's directory board updated; visitors read the board and walk themselves to the office. If the clerk leaves, the board still works — until someone moves offices.

saying these in an interview costs you the question

  • Saying every Service packet is proxied through the kube-proxy process in userspace
  • Claiming kube-proxy assigns Pod IPs or is the CNI plugin
  • Attributing DNS resolution of Service names to kube-proxy
  • Saying all Services stop working the instant kube-proxy dies
  • Describing kube-proxy as doing HTTP path-based routing or TLS termination

context