skip to content

Compare an encapsulated overlay pod network (for example VXLAN) with a routed pod network (for example one that advertises pod routes over BGP): how does each move a packet between two nodes, and what would make you choose one over the other?

level: seniorimportance: should knowfreq 44%

answer

  1. encap = hide pod packet in node-to-node UDP
  2. routed = advertise pod CIDR, native MTU
  3. VXLAN ≈ 50 bytes overhead → pod MTU 1450
  4. MTU bug: small OK, large/TLS hangs
  5. cross-subnet mode = encapsulate only when needed

basics

~20 s

An overlay wraps each pod packet inside a node-to-node UDP packet, so the underlay never needs to know pod addresses — but it costs header bytes (MTU) and hides traffic. A routed network advertises each node's pod range as real routes, so packets travel unwrapped at native MTU, but the underlay must accept those routes.

solid answer

~60 s

**Overlay (VXLAN/Geneve):** the source node wraps the pod packet in a UDP datagram addressed node-to-node, the destination node unwraps and delivers. The underlay only ever sees node IPs, so pod addressing is completely decoupled from the physical network and works across subnets and clouds unchanged. Costs: ~50 bytes of header on IPv4, so pod MTU must drop (1450 on a 1500 underlay); packet captures show tunnels, not pod flows; underlay firewalls and flow logs cannot see pod identity; some CPU per packet. **Routed (BGP or cloud route tables):** each node's pod CIDR is advertised as a route, so packets leave unencapsulated with the pod IP as source. Native MTU, native performance, and the underlay can see, route, and police pod traffic. Costs: you need the underlay to accept the routes — a BGP peering or a cloud route table with quota limits — and pod address space must not collide with the network. Pick routed where you control the fabric and want performance and visibility; overlay where the underlay is opaque, fixed, or spans subnets.

code

bash · 8 lines
bash
ip link show eth0            # expect mtu 1450 on a VXLAN overlay over 1500

# unfragmented probe: 1422 payload + 28 bytes headers = 1450
ping -M do -s 1422 10.244.2.7   # should succeed
ping -M do -s 1472 10.244.2.7   # fails if the overlay overhead is unaccounted for

# on the node: watch the tunnel rather than the pod flow
tcpdump -ni any 'udp port 4789'

go deeper

for a junior

It is enough to say one wraps pod packets inside node-to-node packets and the other advertises pod routes so packets travel as-is.

for a middle

Add the concrete overhead and MTU consequence, name VXLAN and BGP, and explain why an overlay works across subnets without network changes.

for a senior

Lead with the tradeoff axes and the operational evidence: MTU verification, tunnel-level captures, loss of underlay visibility, route-table quotas, and cross-subnet hybrid modes.

for a principal

Treat it as a fabric decision with organisational consequences — who owns address space, whether the network team must approve every cluster, how encryption and observability layer on, and what the migration path looks like if the choice is wrong.

## The problem both solve The pod network model demands that a packet sent from pod A on node 1 to pod B on node 2 arrive with its addresses intact. The physical network, though, knows nothing about pod addresses — it routes node IPs. There are exactly two families of answers: hide the pod packet inside something the underlay understands (encapsulate), or teach the underlay about pod addresses (route). ## Encapsulated overlays The dominant format is **VXLAN**: the original pod-to-pod IP packet becomes the payload of a UDP datagram (destination port 4789 by convention) sent from node 1's IP to node 2's IP, with a VXLAN header carrying a network identifier. Geneve is the more extensible successor; IP-in-IP is a lighter variant with a smaller header but no UDP port to hash on. How a packet moves: the pod sends to 10.244.2.7; the node's routing table says that range is reachable via the tunnel device; the datapath looks up which node owns that range (learned from the API server or the plugin's own store), builds the outer header, and sends. The receiving node's tunnel device strips the header and delivers into the veth of the destination pod. **Strengths.** The underlay never needs configuration — any L3 path between nodes works, including across subnets, availability zones, VPC peerings and mixed clouds. Pod addressing is entirely yours, so overlapping with corporate ranges is harmless. Adding a node needs no network-team ticket. **Costs.** - **MTU.** Roughly 50 bytes of overhead for VXLAN over IPv4 (more for IPv6 or Geneve options). Pod MTU must be lowered accordingly. Getting this wrong is one of the classic incidents: small requests succeed, large payloads or TLS handshakes hang, because oversized packets are dropped and Path MTU Discovery is blocked somewhere. - **Opacity.** Underlay flow logs, IDS and firewalls see node-to-node UDP, not pod-to-pod flows. Troubleshooting requires decapsulating captures. - **CPU.** Per-packet encap/decap work, partly mitigated by NIC offload — but offload for tunnels is uneven, and checksum/segmentation offload behaviour differs between drivers. - **Load spreading.** Because outer headers are similar, some fabrics hash flows poorly across links; implementations use the outer UDP source port for entropy, which is exactly why UDP-based VXLAN beats plain IP-in-IP here. ## Routed networks Here the pod packet is placed on the wire as-is. Each node owns a pod CIDR and announces it. Announcement mechanisms: - **BGP** peering between nodes and the top-of-rack routers or a route reflector — the on-premises answer, and how Calico's routed mode works. - **Same-L2 shortcut** ("host-gw" style): if all nodes share a subnet, each node simply installs a route for every other node's pod CIDR via that node's IP. No protocol needed, but it stops working the moment nodes span subnets. - **Cloud route tables / VPC-native addressing**: the cloud provider programs pod ranges into the VPC route table, or pods take real VPC addresses from secondary interfaces (AWS VPC CNI, GKE alias IPs). Then pod IPs are first-class in the cloud network. **Strengths.** Native MTU, no per-packet work, straightforward captures, and the underlay can apply its own routing, ECMP, ACLs and flow logging to real pod addresses. With VPC-native addressing, cloud load balancers can target pods directly and cloud security groups apply to pods. **Costs.** You need cooperation from the network: a BGP session, or route-table quota (cloud route tables are typically limited to tens of entries, which caps cluster size unless the CNI uses native addressing instead). Pod address space must be globally non-overlapping with everything routable, which turns IP planning into a real design constraint. VPC-native addressing burns real subnet IPs and is bounded by per-instance interface/address limits. ## Hybrids and the practical middle Most mature plugins support both and can switch per-route: encapsulate **only** when the destination node is in a different subnet ("cross-subnet" mode), and route natively inside a subnet. That keeps the fast path unencapsulated while surviving a multi-zone topology. Some deployments add WireGuard or IPsec encryption on top, which is itself another encapsulation with its own MTU cost. ## How to answer the choice question Decide on four axes: (1) do you control the underlay enough to inject routes; (2) is pod address space scarce or must it be non-overlapping; (3) do you need underlay visibility and enforcement of pod traffic; (4) how much per-packet cost and MTU complexity you can carry. On-premises with a cooperative network team, routed wins. Across heterogeneous or opaque networks, or when you need address independence, overlay wins. Whichever you pick, verify MTU end to end with a large-payload test before declaring it working.

  • An application works for small requests but hangs on large POSTs and some TLS handshakes, only between pods on different nodes. What do you suspect?
    An MTU mismatch on the overlay. The pod interface is advertising an MTU the tunnel cannot carry, so full-size packets are dropped once the encapsulation header is added, while small packets pass. TLS handshakes with large certificate chains are a classic trigger. Confirm with ping -M do at increasing sizes across nodes, then set the pod MTU to the underlay MTU minus the encapsulation overhead.
  • Why might a routed pod network limit how large a cluster can get in a cloud VPC?
    If pod CIDRs are advertised by writing entries into the provider's route table, that table has a hard quota — often around 50 to 100 routes — and one entry per node caps the node count. Providers avoid this with VPC-native addressing, where pods take addresses from the subnet directly, but that trades the route-table limit for subnet address exhaustion and per-instance interface limits.

Overlay is putting a letter inside a second envelope addressed to the branch office: the postal service only reads the outer one. Routed is convincing the postal service to deliver to the inner address directly — faster and traceable, but they must agree to know about those addresses.

saying these in an interview costs you the question

  • Assuming an overlay is always slower in a way that matters, without measuring or mentioning offload
  • Forgetting MTU entirely when describing encapsulation
  • Saying a routed network needs no cooperation from the physical network
  • Treating overlay traffic as encrypted because it is 'tunnelled' — VXLAN is plaintext
  • Believing you must choose one globally, when mature plugins encapsulate only cross-subnet

context