skip to content

Overlay Networks & Addressing

Giving every workload an address other hosts can reach, by wrapping its packets inside host-to-host traffic or by routing its range natively. The address pool runs out long before the CPU does.

on this pageshow

questions

6

When a workload's packet is wrapped inside host-to-host traffic to cross the fleet, what does that wrapper cost?

level: middleimportance: must knowfreq 50%

answer

  1. delivery, not protection
  2. who adds it, who removes it
  3. what the fabric can still see
  4. bytes taken out of every packet
  5. per-packet work on both hosts

basics

~20 s

Wrapping costs bytes and work: an outer host-to-host header rides on every packet and leaves fewer usable bytes inside it, both hosts spend processor time adding and stripping it, and the network between them sees only host addresses.

solid answer

~50 s

In an encapsulated design the sending host takes the packet the workload emitted, puts an outer header on it addressed from its own host address to the destination host's, and sends that; the receiving host strips the outer header and delivers the original inward. The fabric in between therefore needs to know nothing about workload addresses, which is the whole appeal - it works on a network you do not control. What you pay is three things. Tens of bytes of every packet are now header rather than data, so the workload's interface must be told to emit smaller packets or large transfers break. Both hosts do per-packet work adding and removing the wrapper, which shows up at high packet rates. And anything that wants to see or act on workload addresses has to run on the hosts, because the path between them carries only host-to-host traffic.

code

pseudocode · 12 lines
pseudocode
# what leaves host A's physical interface
outer   src = host_a_address       dst = host_b_address
wrap    network_id = telemetry_network
inner   src = workload_x_address   dst = workload_y_address
data    ... application bytes ...

# on host B
strip(outer, wrap)
deliver(inner + data) -> workload_y's private network view

# what the network between the hosts ever sees:
#   host_a_address -> host_b_address, and nothing about workloads

go deeper

for a junior

Recall the shape: the sending host puts an outer header addressed host-to-host in front of the workload's packet, and the receiving host takes it off again before delivering it.

for a middle

Explain all three costs - fewer usable bytes per packet, a reduced packet size the workload interface must be told about, and per-packet work on both hosts - and why the design is chosen anyway.

for a senior

Show you know what goes dark: the fabric sees only host addresses, so counting, filtering and balancing on workload identity has to move onto the hosts, and the size headroom becomes something you must plan rather than notice.

for a principal

Frame it as a control decision. Encapsulation buys independence from whoever owns the network and lets address pools repeat between clusters; native routing buys visibility and headroom at the price of real addresses and fabric agreement.

## What the wrapper actually is A telemetry ingester replica on host A sends a packet to a replica on host B. The packet the workload emitted is addressed from one workload address to another - and the network between the two hosts has never heard of either. So the sending host does not put that packet on the wire as it stands. It puts a second header in front of it, addressed from **host A's own address to host B's**, and sends *that*. Host B recognises what arrives, removes the outer header, and delivers the original packet into the destination workload's private network view, which receives it exactly as sent. That is the whole mechanism, and its one property that matters is what it asks of the network in between: **nothing**. The fabric sees ordinary host-to-host traffic, routes it the way it routes everything else, and never needs a route to a workload address. On a network you do not own - a rented estate, a provider's fabric, a site whose network team has no interest in your address plan - this is often the only design available. ## The three costs - **Bytes per packet.** The outer header is on the order of tens of bytes on every single packet. That is bandwidth you paid for and did not use for data, and at high packet rates - which is what thousands of short-lived telemetry replicas produce - it is not a rounding error. - **Size headroom, which is the one that bites.** The link carries a frame of a fixed maximum size. If the workload still believes it may emit packets that fill the link exactly, then wrapping pushes the result *over* the limit and it is dropped. The workload's interface must be told a smaller maximum transmission unit - the link's limit minus the wrapper - or large transfers fail in a way nothing reports. - **Per-packet work on both hosts.** Someone adds the header and someone removes it, for every packet in both directions. Where this work happens varies between designs; that it happens does not. ## What disappears from the fabric's view | | what it sees | where the work must happen | |---|---|---| | the network between hosts | host addresses only | routing, as usual | | the receiving host | both outer and inner | unwrapping | | the receiving workload | the caller's own address | unchanged from the flat model | This is a genuine trade, not just a cost. Because the fabric cannot distinguish one workload's traffic from another's, anything that filters, counts or balances on workload identity has to run on the hosts themselves. Some teams consider that an advantage - the enforcement point moves next to the workload and stops depending on the network team. Others find that the traffic picture they used to get from the fabric goes dark, and have to rebuild it from the hosts. Note what the wrapper does **not** do: it is a delivery header, not protection. It does not encrypt the packet it carries, and it does not hide the caller's address from the receiving workload - the inner header arrives intact. Confidentiality between hosts is a separate, separately-priced decision. ## Why designs accept it anyway Besides working without the network's cooperation, encapsulation buys one more thing: **the pool no longer has to be globally unique**. Because no router outside the hosts ever sees a workload address, two independent clusters can carve their workload addresses out of the very same private pool without conflict. That is a large saving where addresses are scarce, and it is why the alternative - handing workload ranges to the fabric and routing to them natively - costs real addresses from the site's own space. ## The alternative, stated fairly The other design routes workload addresses natively: each host's range is known to the fabric, and a workload packet crosses the network as itself, with no outer header. That returns the bytes, the size headroom and the fabric's view of workload addresses. What it requires is that someone can actually put those routes into the network, and that the address budget can afford ranges that must be unique across the routed domain. Which of the two is right is a genuine judgment call, and it is decided by who controls the fabric far more often than by benchmarks.

  • Why does an encapsulated design let two separate clusters use the very same address pool?
    Because no router outside the hosts ever sees a workload address - only host addresses have to be unique on the fabric. Each cluster can therefore carve its workloads out of the same private pool with no conflict. The catch appears the moment workloads in the two clusters must address each other directly: the overlapping pools then need translation between them, or one of them has to be re-planned.
  • Does wrapping a packet protect it in any way?
    No. The wrapper is a delivery header, not protection: it does not encrypt what it carries, and the caller's own address is still there in the inner header for the receiving workload to see. Some attachment drivers can encrypt host-to-host traffic as a separate feature, at extra per-packet cost. Assuming an overlay is private because it is an overlay is a common and expensive mistake.

It is a second envelope. The inner letter is addressed to a person, the outer one only to the building, and the postal service reads only the outer. The inner letter has to be a little shorter to fit.

saying these in an interview costs you the question

  • Thinks the fabric routes on workload addresses in an encapsulated design.
  • Says wrapping costs bandwidth but never costs packet size.
  • Assumes the network removes the wrapper rather than the receiving host.
  • Treats per-packet wrap and unwrap work as negligible at high packet rates.
  • Believes the wrapper encrypts or hides the traffic it carries.
open as a page

Why do container platforms give every workload its own fleet-routable address instead of sharing each host's address?

level: middleimportance: must knowfreq 62%

basics

~20 s

Giving each workload its own fleet-routable address lets any workload dial any other directly, with no port bookkeeping and no rewritten source address, so callers reach a workload rather than the host it happens to sit on.

open as a page

After a workload overlay is introduced, small requests succeed but large responses hang between hosts - why?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Almost certainly a packet-size mismatch. The workload interface still advertises the link's full size, so once the host adds its wrapper the frame exceeds what the link carries and is dropped. Small exchanges never fill a packet, so they survive.

open as a page

A host with idle processors accepts no more workloads because its address range is full - why, and what fixes it?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Addresses are a capacity dimension of their own. One pool is carved into fixed per-host ranges, so a host whose range is used up places nothing more, whatever its spare processor. The fix is re-planning the ranges or enlarging the pool.

open as a page

How would you choose between an encapsulated overlay and a natively routed workload network for a large estate?

level: principalimportance: should knowfreq 34%

basics

~20 s

Decide on control of the network, the address budget and what you need to see - not on benchmarks. An overlay needs no cooperation and lets pools repeat per cluster; native routing costs real addresses and agreement, and returns headroom and workload-level visibility.

open as a page

When a platform starts a workload, what creates its network interface and gives it an address?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

Not the platform itself. It defines a small contract and calls a pluggable attachment driver before the workload's first process starts, handing it the workload's private network view; the driver builds the interface, takes an address from the host's range and returns it.

open as a page