skip to content

Kubernetes' kube-proxy component can run in iptables mode or IPVS mode. Compare how each implements Service load balancing, and explain what actually changes as the number of Services and endpoints grows.

level: middleimportance: must knowfreq 62%

answer

  1. iptables: chains + random probability, linear walk
  2. IPVS: hash tables, incremental updates, real schedulers
  3. IPVS still needs ipset plus some iptables
  4. Sync latency is the real pain, not per-packet cost
  5. nftables mode = verdict maps, GA in 1.33

basics

~20 s

iptables mode builds a linear chain of NAT rules per Service and picks a backend with random-probability jumps; rule count grows with Services times endpoints, and every change means re-syncing a large ruleset. IPVS mode uses the kernel's L4 load balancer with hash-table lookups, so matching is near constant time, updates are incremental, and real schedulers (round-robin, least-connection, source-hash) are available.

solid answer

~60 s

**iptables mode** translates each Service into chains in the `nat` table: a dispatch chain matching the virtual IP and port, then a chain per endpoint selected by `statistic mode random probability`. Matching walks rules sequentially, so cost is roughly linear in rules, and rule count scales with Services times endpoints. Historically the painful part was update latency - kube-proxy rewrote large rulesets on every change, so convergence in clusters with thousands of Services could take seconds (modern versions do partial syncs, which helps a lot). **IPVS mode** uses the in-kernel L4 load balancer built for exactly this. Each Service becomes a virtual server with real servers behind it, looked up via hash tables, so lookup cost is effectively constant and adding an endpoint is an incremental update rather than a ruleset rewrite. It also exposes scheduling algorithms - `rr`, `lc` (least connection), `sh` (source hashing) and others - instead of iptables' flat random split. IPVS still needs some iptables and ipset rules for masquerading and NodePort handling, so it is not a full replacement. Newer clusters increasingly use **nftables mode** instead, which keeps the iptables programming model but with map-based lookups and far better update behaviour.

code

bash · 8 lines
bash
kubectl -n kube-system get cm kube-proxy -o yaml | grep -i mode
curl -s localhost:10249/proxyMode   # on a node

# iptables mode
sudo iptables -t nat -S | grep KUBE-SVC | head

# IPVS mode
sudo ipvsadm -Ln

go deeper

for a junior

Know that kube-proxy has more than one backend mode, that iptables is the common default, and that IPVS exists for larger clusters.

for a middle

Explain the actual mechanisms: random-probability chains versus IPVS virtual and real servers, and why rule count and sync time scale differently.

for a senior

Talk about the operational symptom that drives the switch - endpoint convergence lag after deploys - and the prerequisites and debuggability tradeoffs of IPVS.

for a principal

Position mode choice against the wider dataplane strategy: nftables as the forward path, eBPF kube-proxy replacement as a way to drop conntrack and DNAT cost, and the coupling that creates with CNI selection.

## The same contract, two implementations Both modes implement identical Service semantics: a virtual IP is DNATed to one of the ready backend pods, chosen per connection. What differs is the kernel machinery, and therefore the scaling curve. ## iptables mode kube-proxy generates a hierarchy of chains in the `nat` table: - `KUBE-SERVICES` is entered from PREROUTING/OUTPUT and contains one match per Service VIP and port. - Matching jumps to a per-service chain (`KUBE-SVC-XXXX`), which contains one rule per endpoint. - Each endpoint rule uses `-m statistic --mode random --probability p` with probabilities computed so the split is even (1/n, then 1/(n-1) of the remainder, and so on). The final rule is unconditional. - The endpoint chain (`KUBE-SEP-XXXX`) does the actual `DNAT --to-destination podIP:port`, plus masquerade marking where needed. Two costs follow. First, **matching is a sequential walk**: netfilter evaluates rules in order, so per-packet cost grows with the number of Services - but only for the first packet of a flow, since conntrack short-circuits the rest, which is why the practical impact is smaller than the theory suggests. Second, and historically worse, **update cost**: iptables has no incremental API, so kube-proxy rendered the full ruleset and pushed it through `iptables-restore`. In a cluster with tens of thousands of endpoints, a single pod churn could trigger a multi-second rewrite while holding the xtables lock. Modern kube-proxy mitigates this with partial syncs, `minSyncPeriod` batching, and only rewriting changed chains. ## IPVS mode IPVS (IP Virtual Server) is a mature kernel L4 load balancer that predates Kubernetes. In this mode kube-proxy: - Creates a **virtual server** per Service VIP and port and a **real server** per endpoint, using hash tables, so backend lookup is effectively constant time regardless of Service count. - Applies updates **incrementally** - adding or removing one real server touches one entry, not the whole ruleset. Convergence stays flat as the cluster grows. - Supports real **schedulers**: `rr` (round robin, the default), `wrr`, `lc` (least connection), `wlc`, `sh` (source hashing, useful for affinity), `dh`, `sed`, `nq`. iptables mode offers only an even random split. - Binds Service VIPs to a dummy interface (`kube-ipvs0`), which makes the addresses visible locally - a small behavioural difference that occasionally surprises people. IPVS is not iptables-free. kube-proxy still uses **ipset** plus a small set of iptables rules for masquerading, NodePort, loadBalancerSourceRanges, and reject-when-no-endpoints. Using ipset keeps that ruleset roughly constant in size instead of growing per Service. The node also needs the `ip_vs`, `ip_vs_rr`, `ip_vs_wrr`, `ip_vs_sh` and `nf_conntrack` modules loaded, which is a classic reason IPVS mode silently falls back to iptables on a misconfigured node. ## nftables mode The newer nftables mode is the strategic direction on Linux: it keeps kube-proxy's programming model but uses nftables **verdict maps**, so Service lookup is a hash lookup rather than a linear scan, and updates are incremental transactions. For new clusters on recent kernels it removes most of iptables mode's scaling problems without IPVS' extra kernel-module dependency. ## Choosing between them in practice - **Small to medium clusters** (hundreds of Services): all modes behave fine; iptables mode's simplicity and ubiquity win. - **Large clusters** (thousands of Services, tens of thousands of endpoints): iptables mode's sync latency shows up as slow convergence after deploys - traffic to just-deleted pods, or new pods not receiving traffic for seconds. IPVS or nftables fixes that. - **Needing a scheduler other than even random** (least-connection for long-lived uneven work, source hashing for stickiness): only IPVS offers it. - **Debuggability**: iptables mode is inspectable with tooling everyone knows; IPVS needs `ipvsadm`, which fewer engineers have used. An increasingly common fourth option is deleting kube-proxy entirely and letting an eBPF-based CNI implement Service semantics in the socket and TC layers, which removes both DNAT-on-every-connection and conntrack pressure - but that couples your Service dataplane to your CNI choice.

  • Why is per-packet cost in iptables mode less damaging in practice than the linear rule-walk suggests?
    Because only the first packet of a connection traverses the NAT chains. Once conntrack has an entry, subsequent packets of that flow are matched and translated by the connection-tracking fast path without re-walking the ruleset. The dominant real-world cost is therefore control-plane sync time when endpoints change, not steady-state forwarding throughput.
  • You switch a large cluster to IPVS mode and Services stop working on a subset of nodes. What do you check first?
    Whether the required kernel modules (ip_vs, ip_vs_rr, ip_vs_wrr, ip_vs_sh, nf_conntrack) are loaded on those nodes, and what kube-proxy logged at startup - it warns and can fall back to iptables when IPVS is unavailable. Also confirm ipset is installed, since IPVS mode relies on it for masquerade and NodePort rules.

saying these in an interview costs you the question

  • Saying IPVS makes load balancing L7 or request-aware - it is still L4 and per-connection
  • Claiming IPVS mode eliminates iptables entirely from the node
  • Assuming iptables mode's cost is mainly per-packet throughput rather than ruleset sync latency
  • Thinking switching modes changes Service semantics or requires application changes
  • Believing you can choose an IPVS scheduler per Service - it is a kube-proxy-wide setting

context