Give an overview of Flannel, Calico and Cilium as Kubernetes pod network implementations: what each is built on, and what would push you toward one rather than another?
answer
- Flannel = minimal VXLAN, no policy
- Calico = BGP routed or VXLAN, iptables or eBPF, flexible IPAM
- Cilium = eBPF datapath, identity policy, kube-proxy replacement, Hubble
- cross-subnet mode = encapsulate only when needed
- cloud CNIs = real VPC IPs, address-limit tradeoff
basics
~20 sFlannel is the minimal option — usually a VXLAN overlay, simple, no policy enforcement. Calico offers routed (BGP) or encapsulated modes with a mature policy engine on iptables or eBPF. Cilium builds its datapath on eBPF, adds identity-aware and layer-7 aware policy, can replace kube-proxy, and brings its own flow observability.
solid answer
~50 s**Flannel** does one thing: give pods a flat network. Typically VXLAN, sometimes a same-subnet host-gw route mode; IPAM is a simple per-node subnet. It has no policy engine of its own, so isolation needs a second component. Choose it for small or learning clusters where simplicity is the point. **Calico** is the general-purpose choice: routed mode with BGP (peering to real routers or via route reflectors), or VXLAN/IP-in-IP where you cannot inject routes, including cross-subnet-only encapsulation. Its dataplane can be iptables/nftables or eBPF. Flexible IPAM, mature policy enforcement, strong on-premises story. **Cilium** puts an eBPF program in the kernel datapath instead of long iptables chains. That gives efficient service handling (it can replace kube-proxy), policy keyed on workload identity rather than IP, awareness of application protocols, per-flow visibility via Hubble, and options like WireGuard encryption and multi-cluster meshing. The cost is a higher kernel-version floor and more concepts to operate. Managed clusters also ship provider CNIs that give pods real VPC addresses.
go deeper
One clear sentence each is enough: Flannel simple overlay, Calico routed or overlay with policy, Cilium eBPF-based with richer features.
Add the datapath mechanics and the enforcement capability, and note the kernel requirement that comes with eBPF.
Drive the comparison from constraints — underlay control, address space, service count, observability, encryption — and mention cross-subnet hybrids and provider plugins.
Frame plugin choice as a long-lived, hard-to-reverse platform decision: it dictates IP planning, upgrade risk (a DaemonSet rollout touches the datapath), the observability toolchain, and which enforcement model your teams will learn.
## How to compare implementations at all Every pod network plugin answers four questions, and comparing them is easiest along those lines: 1. **Datapath** — how does a packet get from pod to pod (overlay, routed, eBPF-accelerated)? 2. **IPAM** — where do pod addresses come from and how flexibly? 3. **Enforcement** — can it enforce policy, and with what mechanism? (Policy *semantics* are a separate subject; here only capability matters.) 4. **Operability** — kernel requirements, observability, upgrade story, ecosystem features like encryption or multi-cluster. ## Flannel The deliberately small option, from the CoreOS lineage. A per-node agent assigns each node a subnet from the cluster CIDR and programs a datapath — by default `vxlan`, optionally `host-gw` (plain routes when all nodes share a subnet) or provider-specific backends. Configuration is a single ConfigMap; there is little to tune and little to break. What it does not do: enforce network policy. Historically deployments paired it with a policy component (the "Canal" combination) to fill the gap. It also offers no fancy IPAM, no encryption, no built-in visibility. **Choose it when** the cluster is small, you want the fewest moving parts, and isolation and observability are handled elsewhere or not yet needed. ## Calico The broadly deployed general-purpose plugin, and the one that popularised routed pod networking. Its node agent (Felix) programs the node's datapath while a routing component distributes pod routes. - **Datapath options**: native routing with **BGP** (peer directly with top-of-rack switches, or use route reflectors so you do not need a full mesh at scale), or encapsulation with VXLAN or IP-in-IP. Crucially it supports *cross-subnet* mode: encapsulate only when source and destination nodes are in different subnets, route natively otherwise. - **Dataplane**: the traditional iptables/nftables implementation, or an eBPF dataplane for lower overhead and better service handling. - **IPAM**: multiple pools, per-namespace or per-node pool selection, block sizes you control — useful when address space is constrained or when certain workloads need routable addresses. - **Enforcement**: a mature policy engine, including its own richer policy resources beyond the core Kubernetes objects. **Choose it when** you run on-premises with a network you can peer with, when you need address-space flexibility, or when you want a well-understood default with strong policy support and a long operational track record. ## Cilium Cilium's premise is that the kernel's programmable **eBPF** hooks are a better place to build the datapath than long chains of iptables rules. Programs attached at the network interface and socket layers handle forwarding, address translation and policy directly. What that unlocks: - **kube-proxy replacement**: Service virtual IPs are handled by eBPF programs (often at the socket layer for local traffic), which scales better than rule chains that grow with the number of services and backends. - **Identity-based enforcement**: workloads get an identity derived from their labels, and policy is evaluated on identity rather than on ephemeral IPs — which suits churn-heavy clusters. - **Protocol awareness**: policy can consider application-layer attributes such as HTTP paths or Kafka topics, not just address and port. - **Observability**: Hubble exposes per-flow records with pod-level identity, which is a real advantage on an encapsulated datapath where captures would otherwise show only tunnels. - **Extras**: transparent WireGuard or IPsec encryption, multi-cluster meshing, egress gateway, bandwidth management. Costs: a modern kernel is required, and the more features you turn on the more concepts and control-plane components you operate. Debugging shifts from `iptables-save` to Cilium's own tooling. **Choose it when** you have large service counts, need identity or protocol-aware enforcement, want built-in flow observability or encryption, and can guarantee the kernel baseline. ## Managed and provider plugins On managed clusters, the default is often a provider plugin that gives pods **real VPC addresses** from secondary interfaces or alias ranges. Then there is no overlay at all: cloud load balancers can target pods directly, cloud security groups and flow logs apply to pods, and latency is native. The tradeoff is address-space pressure — pods consume real subnet IPs — plus per-instance limits on how many addresses a node can hold, which caps pod density. Several providers now let you run Cilium or Calico in place of, or layered on, the default. ## Answering well Avoid ranking them absolutely. State each one's datapath and enforcement story in a sentence, then choose against the constraints of the environment described: underlay control, address space, kernel version, service count, observability and encryption needs, and how much operational surface the team can carry.
- What does 'replacing kube-proxy' actually buy you with an eBPF-based plugin?Service handling moves from rule chains that grow with the number of services and endpoints into eBPF programs with map lookups, so lookup cost stays roughly flat as the cluster grows and rule-programming latency after endpoint churn drops. Some implementations translate at the socket layer for pod-local connections, avoiding per-packet translation entirely. The tradeoff is a kernel-version floor and a different set of debugging tools.
- Why might an organisation deliberately keep an overlay even though its network team could support routed pod networking?Address independence and operational autonomy. With an overlay, pod ranges never have to be coordinated with corporate IP planning, they can overlap between clusters, and adding nodes or clusters requires no change on the physical network. The team accepts MTU management and reduced underlay visibility in exchange for not needing a network change ticket for every cluster event.
saying these in an interview costs you the question
- Declaring one plugin universally 'the best' without naming a constraint
- Saying Flannel enforces network policy on its own
- Believing eBPF is only a performance feature, missing identity-based enforcement and observability
- Assuming Calico is always BGP-routed — it also has VXLAN and cross-subnet modes
- Thinking cloud provider plugins avoid all limits, ignoring per-node address limits and subnet exhaustion