skip to content

You run a multi-tenant Kubernetes platform heading toward tens of thousands of Services and very high pod churn. How would you decide between keeping kube-proxy in iptables mode, moving to IPVS or nftables, or removing kube-proxy in favour of an eBPF dataplane?

level: principalimportance: nice to knowfreq 32%

answer

  1. Same semantics, different coupling and failure modes
  2. Trigger on sync-duration plus convergence SLO
  3. nftables = lowest-risk in-family step
  4. eBPF = flat curve, vendor coupling, kernel floor
  5. Fix drain windows before blaming the proxy

basics

~20 s

Decide on measured convergence and failure behaviour, not benchmarks. Keep iptables while endpoint-sync latency stays inside your rollout error budget; move to nftables or IPVS when sync time is the bottleneck; drop kube-proxy for eBPF only when you accept coupling the Service dataplane to one CNI and can staff kernel-level debugging.

solid answer

~60 s

I would treat this as a **coupling and operability** decision framed by measured signals. First, instrument: kube-proxy sync duration, endpoint propagation time end to end, conntrack table utilisation, and the user-visible metric that actually matters - error rate during rollouts. Those tell you whether you have a dataplane problem or an application drain problem. Then the ladder. **iptables mode** is the default everywhere, universally debuggable, and fine until sync latency during large rollouts eats the error budget. **nftables mode** (GA in v1.33) keeps kube-proxy's model and semantics but replaces linear chains with map lookups and incremental updates - the lowest-risk step because nothing above it changes. **IPVS** buys hash lookups plus real schedulers like least-connection, at the cost of extra kernel modules and `ipvsadm`-shaped debugging that fewer engineers have. **Removing kube-proxy** for an eBPF CNI removes DNAT and much conntrack pressure and gives per-Service scale the others cannot reach - but it welds Service semantics to one CNI vendor, raises the kernel-version floor, and moves incident debugging into eBPF tooling. On a multi-tenant platform I would want that decision backed by a real capacity forecast and a staffed exit path.

go deeper

for a junior

Know that kube-proxy has multiple modes and that some CNIs can replace it entirely; you are not expected to drive this choice.

for a middle

Be able to compare the modes mechanically and name the prerequisites - kernel modules for IPVS, kernel and version floor for nftables.

for a senior

Lead with measurement: which metrics prove the dataplane is the bottleneck, and how you roll a change per node pool with a revert path.

for a principal

Own the coupling and staffing argument - forecast scale, weigh vendor lock-in against a flat scaling curve, and set a convergence SLO that triggers the decision rather than deciding by preference.

## Frame the decision, not the benchmark All four options implement the same Service contract, so the question is never "which is faster" in isolation. It is: which failure modes do we accept, what do we couple ourselves to, and who debugs it at 3am. Benchmarks comparing packets-per-second mislead, because in real clusters the first packet of a flow pays the rule-walk and every subsequent packet rides conntrack - steady-state throughput differences are usually small, while **convergence latency** differences are large. ## Signals that should drive the move Before changing anything, define the trigger with data: - **kube-proxy sync duration** (`kubeproxy_sync_proxy_rules_duration_seconds`) and last-sync timestamp per node. Rising p99 during deploys is the canonical iptables-mode smell. - **End-to-end endpoint propagation**: time from pod ready to traffic actually arriving, measured synthetically. - **Error budget consumption during rollouts** - the only signal tenants feel. - **conntrack utilisation** (`nf_conntrack_count` against max). Table exhaustion produces random drops that look like application bugs and is a distinct failure mode from rule-count scaling. - **Object counts**: Services, EndpointSlices, endpoints per Service, churn rate. These give you the forecast curve rather than a snapshot. ## The options and what each really costs **Stay on iptables.** Ubiquitous, every engineer can read the rules, every vendor supports it. Modern kube-proxy does partial syncs and batches with `minSyncPeriod`, so the historical cliff is softer than its reputation. Cost: sync time still grows with ruleset size, and at very high Service counts convergence lag becomes visible as traffic to dead pods and slow-to-receive new pods. **Move to nftables mode.** The smallest-blast-radius step: same component, same semantics, same operational model, but verdict maps instead of linear chains and transactional incremental updates. It needs a recent kernel and a version where the mode is GA (v1.33), and your node tooling and any third-party agents that write iptables rules must coexist cleanly - that coexistence, not kube-proxy itself, is usually where trouble appears. **Move to IPVS.** Hash-based lookup and incremental real-server updates, plus schedulers (`lc`, `sh`, `wrr`) that iptables cannot express - genuinely useful when backends have uneven, long-lived work. Cost: extra kernel modules that must be present on every node image (a silent fallback to iptables if not), a residual ipset and iptables layer that still needs understanding, and a debugging surface fewer people know. Increasingly it looks like a transitional answer now that nftables mode exists. **Remove kube-proxy entirely.** An eBPF CNI implementing Service semantics at the socket layer can rewrite the destination at connect() time, so there is no per-packet DNAT and much less conntrack dependence, with lookups in eBPF maps that scale flat. This is the only option that changes the shape of the curve rather than its constant. The costs are strategic: Service load balancing, NetworkPolicy enforcement, and observability all become one vendor's concern; the kernel-version floor rises and node upgrades become a coupled dependency; incident tooling shifts to eBPF-specific tracing; and reverting means a cluster-wide dataplane migration, not a config flag. ## How I would sequence it on a multi-tenant platform 1. **Instrument and set an SLO** on endpoint convergence, so the change is triggered by budget burn rather than by architecture taste. 2. **Exhaust cheap wins first** - a platform-wide pod template with a preStop drain window fixes most rollout errors regardless of proxy mode, and topology-aware routing cuts cross-zone traffic without touching the dataplane. If teams blame kube-proxy for what is really missing drain handling, changing modes buys nothing. 3. **Take the in-family step** (nftables, or IPVS if you need its schedulers) when sync latency is genuinely the bottleneck. Roll it per node pool, compare the same metrics, keep the ability to flip back with a DaemonSet config change. 4. **Treat kube-proxy replacement as an architecture decision**, not an optimisation: forecast Service counts 18 months out, prototype on a non-critical pool, verify NetworkPolicy and observability parity, and get explicit agreement on the vendor coupling and the staffing needed to debug it. ## Multi-tenancy specifics Tenant isolation raises the stakes: a shared dataplane means one tenant's churn degrades everyone's convergence, so per-tenant Service and churn quotas matter as much as the proxy choice. Blast radius also argues for node-pool-scoped rollouts and for keeping at least one pool on the boring option during any migration.

  • What would make you reject the eBPF kube-proxy-replacement option even at very large scale?
    If the organisation cannot absorb the coupling: a regulated environment requiring multiple CNI vendors, node images pinned to older kernels, or an on-call rota with no eBPF debugging skill. The dataplane must be operable by the people who own incidents. I would also reject it if measurement showed the pain was really application drain behaviour or conntrack sizing, since neither is fixed by swapping the proxy.
  • How do you migrate proxy modes without a cluster-wide outage?
    Roll it per node pool. kube-proxy's mode is per-node configuration, and Service semantics are identical across modes, so nodes in different modes coexist. Cordon and drain one canary pool, switch its kube-proxy config, confirm sync duration and synthetic convergence improve with no regression in tenant error rates, then expand. Keep the previous configuration one flag away, and never migrate during a freeze or a large tenant rollout.

saying these in an interview costs you the question

  • Choosing based on packets-per-second benchmarks rather than convergence latency and failure modes
  • Treating a kube-proxy replacement as a drop-in optimisation instead of a CNI coupling decision
  • Ignoring conntrack table sizing, which causes scale failures unrelated to proxy mode
  • Assuming a mode change fixes rollout errors that are actually missing preStop drain windows
  • Proposing a cluster-wide flag day instead of a per-node-pool rollout with a revert path

context