skip to content

Why does a VXLAN VTEP derive the outer UDP source port from the inner packet's headers, and what happens in a leaf-spine underlay when every flow gets one port?

level: seniorimportance: should knowfreq 22%

answer

  1. what the underlay can actually see
  2. four of five fields never change
  3. one flow, one port, one path
  4. dynamic range from 49152

basics

~20 s

Between two VTEPs every outer field except the UDP source port is fixed, so a per-flow hash carried in that port is the underlay ECMP's only entropy. With one port for every flow, each VTEP pair's traffic rides a single path.

solid answer

~50 s

Underlay routers balance traffic by hashing outer headers, and between two VTEPs the outer source and destination addresses, protocol 17 and destination port 4789 never change. RFC 7348 therefore recommends computing the **source port as a hash of fields from the inner packet**, RECOMMENDED to fall in the dynamic range `49152-65535`, to give ECMP 'a level of entropy'. Each inner flow maps to one port, so it stays on one path and its packets stay in order, while different flows spread across the spines. If every flow gets the same port, the underlay sees one flow per VTEP pair: all of it hashes onto one spine while the others carry none of it, one link congests, and losing that link moves everything at once. The same collapse happens if the underlay hashes only IP addresses, or if the VTEP hashes only inner MAC addresses - all traffic through one gateway MAC then shares a port.

go deeper

for a junior

Recall that VXLAN puts a changing number in the outer UDP source port so the network underneath can spread different conversations over different links.

for a middle

Explain which outer fields are fixed between two VTEPs, why the source port is the one left to vary, and why one inner flow keeps one port.

for a senior

Diagnose lopsided spine links: fixed or MAC-only source ports, an underlay hashing on addresses alone, elephant flows, and what a link failure rehashes.

for a principal

Weigh balance against ordering and hardware: per-flow ECMP caps a single flow, ECMP width must match the spines, and consistent hashing costs table space.

## What an underlay router can see In a **leaf-spine** underlay every leaf reaches every other leaf over several equal-cost paths, one per spine. Routers split traffic across such paths with **ECMP** (equal-cost multipath), which works per flow: the router hashes header fields that identify a flow and maps the result to one next hop. RFC 2992 describes the classic *hash-threshold* method, for example a CRC16 over the flow's fields, with each next hop owning a region of the hash space. A **VXLAN** packet (RFC 7348) shows the underlay only its outer headers. Between two VTEPs (VXLAN Tunnel End Points), A and B: | Outer field | Value between VTEP A and VTEP B | |---|---| | Source IP address | A's tunnel address - fixed | | Destination IP address | B's tunnel address - fixed | | IP protocol | 17, UDP - fixed | | UDP destination port | 4789, IANA's VXLAN port - fixed | | UDP source port | chosen by the sending VTEP - free to vary | The Geneve specification, RFC 8926, names the general problem: encapsulated traffic is hidden from the fabric by design, and without help only the tunnel endpoint addresses are available for hashing. ## The source port as an entropy field RFC 7348 section 5 answers it: the UDP source port is recommended to be **a hash of fields from the inner packet** - its example is the inner Ethernet frame's headers - "to enable a level of entropy for the ECMP/load-balancing" of the tenant traffic. When the port is computed this way the RFC RECOMMENDS keeping it in the **dynamic/private range 49152-65535** (RFC 6335), which gives 16,384 possible values. Geneve, by contrast, allows the entire 16-bit range and notes that an IPv6 underlay MAY also carry entropy in the flow label. ```pseudocode // sending VTEP, for each encapsulated frame key = hash(inner.src_ip, inner.dst_ip, inner.protocol, inner.src_port, inner.dst_port) outer.udp.src_port = 49152 + (key mod 16384) // stays within 49152-65535 outer.udp.dst_port = 4789 // underlay router, for each packet flow = (outer.src_ip, outer.dst_ip, 17, outer.udp.src_port, 4789) next_hop = ecmp_set[hash(flow) mod size(ecmp_set)] ``` ## Per flow, not per packet - **One inner flow, one port.** The same inner headers always produce the same source port, so every packet of a tenant TCP connection follows the same underlay path and arrives in order. - **Different flows, different ports.** Different inner flows usually get different ports, so the underlay spreads them across all the spines. - **No per-packet spraying.** A fresh random port per packet would balance more evenly but reorder each flow across paths with different queues, and TCP reads reordering as possible loss. - **One flow stays on one link.** A single very large tenant flow is capped by one path's capacity; per-flow ECMP cannot split it, by design. ## When the entropy disappears | Cause | What the underlay sees | Result | |---|---|---| | The VTEP uses one fixed source port | One flow per VTEP pair | All of a pair's traffic on one spine, the rest idle for that pair | | The VTEP hashes only inner MAC addresses | One port per MAC pair | Every flow between two hosts, or through one gateway MAC, shares a path | | The underlay hashes only IP addresses (an implementation choice) | Source port ignored | Same collapse as a fixed port | | One elephant tenant flow | One port | One link carries it, as designed | | ECMP stages stacked with little fresh entropy | Pre-sorted flows at the later stage | Uneven link use - the flow polarization RFC 7938 warns of | ## What to check in a design 1. The VTEPs derive the port from inner IP addresses and ports where the inner frame carries IP, not from MAC addresses alone. 2. The underlay's hash includes layer 4 ports for UDP - which fields a device hashes is an implementation choice. 3. The ECMP width on each leaf covers every spine. 4. A link failure rehashes flows; consistent hashing (RFC 7938 section 6.4, RFC 2992) limits how many established flows move, at the cost of more hardware table space.

  • Why not choose a random source port for every packet to spread load perfectly?
    Because ECMP would then send packets of one tenant TCP connection down different paths with different queues, and they would arrive out of order. TCP reads reordering as possible loss, sending duplicate ACKs and needless retransmissions. One port per inner flow trades perfect balance for in-order delivery, which is why both VXLAN and Geneve tie the port to a flow hash.
  • Can good source-port entropy spread one very large tenant flow over all the spines?
    No. One inner flow hashes to one source port and so to one path at each hop; that is what keeps it in order. A single elephant flow is capped by one path's capacity, and two elephants can still collide on one link by chance. Entropy spreads many flows evenly; it never splits one.

A toll plaza that assigns each car to a lane by its licence plate keeps every car from one trip in one lane and spreads different cars across lanes. If a depot fits the same plate on every lorry, the whole fleet queues in one lane while the others stand empty. The UDP source port is the plate a VTEP prints for each inner flow.

saying these in an interview costs you the question

  • Underlay routers must parse the VXLAN payload to balance tenant flows.
  • A random source port per packet balances best and has no side effects.
  • With good source-port entropy, one huge tenant flow is spread across every spine.
  • Hashing only inner MAC addresses gives as much entropy as hashing inner IP addresses and ports.
  • VXLAN varies the UDP destination port so the underlay can balance flows.