skip to content

In VXLAN, how do two VTEPs carry a frame from a VM on rack 1 to a VM on rack 2 in the same segment?

level: middleimportance: must knowfreq 32%

answer

  1. VNI plus destination MAC
  2. lookup gives a remote VTEP address
  3. outer header loopback to loopback
  4. underlay rewrites only outer fields
  5. remote VTEP checks, strips, delivers

basics

~20 s

The source VTEP maps the destination MAC in the VM's VNI to a remote VTEP and wraps the frame in VXLAN, UDP and an outer VTEP-to-VTEP IP header; the underlay routes that packet, and the remote VTEP unwraps it and delivers the unchanged frame.

solid answer

~50 s

VM-A sends an ordinary Ethernet frame to VM-B's MAC address. Its VTEP knows VM-A's port is in VNI `5001`, looks up VM-B's MAC in that VNI and finds remote VTEP `10.0.0.12`. It drops the inner FCS and prepends a VXLAN header with VNI `5001`, UDP to port `4789` with a source port hashed from the inner headers, an outer IP header `10.0.0.11` to `10.0.0.12`, and an outer Ethernet header addressed to its next-hop router. The leaf, spine and leaf route it like any UDP packet, each rewriting only outer fields: the Ethernet header, the TTL and its checksum. VTEP `10.0.0.12` sees UDP `4789` addressed to itself, checks that VNI `5001` is valid and that VM-B is a local host on it, strips the outer headers and delivers the frame. VM-B receives exactly what VM-A sent.

code

pseudocode · 17 lines
pseudocode
encapsulate(frame, port):
    vni = vni_of(port)                              # 5001
    remote = mac_table.lookup(vni, frame.dst_mac)  # 10.0.0.12
    if remote is none: return bum_handling(frame)   # not shown
    inner = frame without FCS
    pkt = vxlan(I=1, vni) + inner
    pkt = udp(src=hash(inner headers), dst=4789) + pkt
    pkt = ip(src=my_vtep_ip, dst=remote, proto=17) + pkt
    route pkt toward remote                         # outer Ethernet per hop

decapsulate(pkt):                                   # unicast case only
    if pkt.ip.dst != my_vtep_ip or pkt.udp.dst != 4789: return
    if pkt.vxlan.I != 1 or pkt.vxlan.vni not served here: drop
    frame = pkt.inner
    port = local_port(pkt.vxlan.vni, frame.dst_mac)
    if port is none: drop
    deliver frame on port

go deeper

for a junior

Remember the shape of the trip: the source VTEP wraps the frame, the underlay routes the wrapper between VTEP addresses, the far VTEP unwraps it.

for a middle

Walk all the steps with real values: the VNI and MAC lookup, the four headers and their key fields, per-hop outer rewrites, then the receiver's checks and delivery.

for a senior

Use the walk to locate faults: a missing table entry, an unreachable VTEP address or a VNI mismatch each break a different step, and the field table tells you where to capture.

for a principal

Use the walk to argue placement: hypervisor VTEPs put the overlay boundary inside every server, top-of-rack VTEPs put it at the leaf, and each choice moves operational ownership.

## The setup Two virtual machines share one VXLAN segment but sit in different racks, and the racks are joined by a routed leaf-and-spine network. Each server runs a software **VTEP** in its hypervisor, and each VTEP's address is a host route the underlay carries. | Item | Rack 1 | Rack 2 | |---|---|---| | Virtual machine | VM-A, `192.168.10.21` | VM-B, `192.168.10.22` | | VM MAC address | `00-00-5E-00-53-01` | `00-00-5E-00-53-02` | | VTEP address (hypervisor) | `10.0.0.11` | `10.0.0.12` | | Segment | VNI `5001` | VNI `5001` | | Underlay path | leaf 1, then a spine | then leaf 2 | VM-A already knows VM-B's MAC address. How VM-A resolved it, and how VTEP `10.0.0.11` came to know that this MAC sits behind `10.0.0.12`, are the subjects of VXLAN's learning or control-plane mechanisms; this walk starts once the mapping exists. ## Step by step 1. **VM-A sends a normal frame**: destination MAC `00-00-5E-00-53-02`, source MAC `00-00-5E-00-53-01`, an IP packet from `192.168.10.21` to `192.168.10.22` inside. VM-A knows nothing of VXLAN. 2. **The source VTEP picks the segment.** VM-A's virtual port belongs to VNI `5001`. 3. **It finds the far end.** Looking up `00-00-5E-00-53-02` within VNI `5001` returns remote VTEP `10.0.0.12`. 4. **It encapsulates.** The original FCS is dropped and four headers are prepended: the VXLAN header with the I flag set and VNI `5001`; UDP with destination port `4789` and a source port that RFC 7348 recommends be a hash of the inner headers; outer IP from `10.0.0.11` to `10.0.0.12`, protocol 17; and an outer Ethernet header whose destination is the MAC address of leaf 1, the next-hop router. A new FCS closes the frame. 5. **The underlay routes it.** Leaf 1, the spine and leaf 2 forward on the outer destination `10.0.0.12`. Each one writes a new outer Ethernet header and decrements the outer TTL, exactly as for any routed packet. 6. **The destination VTEP accepts it.** The packet is UDP to port `4789` addressed to `10.0.0.12`. RFC 7348 has the VTEP verify that the VNI is valid and that a local VM on VNI `5001` uses the inner destination MAC. 7. **It decapsulates and delivers.** The outer headers are stripped and the inner frame goes to VM-B's virtual port. In data-plane learning mode the VTEP also records that `00-00-5E-00-53-01` sits behind `10.0.0.11`, so the reply needs no flooding. The reply makes the same trip in reverse, with `10.0.0.12` as the outer source. ## What changes on the way and what does not | Field | Source VTEP | Underlay routers | Destination VTEP | |---|---|---|---| | Outer Ethernet header | created | rewritten at every hop | removed | | Outer IP TTL | set | decremented at every hop | removed | | Outer IP addresses | set, VTEP to VTEP | unchanged | removed | | VXLAN header and VNI | added | not used for forwarding | checked, removed | | Inner frame and inner TTL | unchanged | unchanged | delivered unchanged | RFC 8014 states the last row as a rule of Layer 2 overlay service: neither the IPv4 TTL nor the IPv6 Hop Limit of the tenant packet is modified, while the underlay manages the TTL of the outer header. ## Where the VTEP sits does not change the walk If the VTEPs live in the top-of-rack switches instead of the hypervisors, the same steps happen one box later. The server sends a plain frame, often VLAN-tagged on its access link; leaf 1 maps it to VNI `5001` and encapsulates with its own VTEP address as the outer source; leaf 2 decapsulates and forwards the frame out of the access port toward VM-B's server. The underlay then begins at the leaf rather than inside the server. ## What the walk leaves to other mechanisms - **Populating the MAC-to-VTEP table**, and what happens when the lookup in step 3 misses, belong to flood-and-learn or to a control plane. - **Frame size**: the added headers make the packet larger than the frame VM-A sent, so the underlay MTU must allow for them. - **Path choice**: when the underlay has several equal-cost paths, the hashed UDP source port is what lets different flows between the same two VTEPs take different spines.

  • In VXLAN, what outer destination MAC address does the encapsulated packet carry as it leaves the source VTEP?
    Usually the MAC address of the next-hop router toward the remote VTEP, not the remote VTEP's own. RFC 7348 says the outer destination MAC may be the target VTEP's or an intermediate Layer 3 router's; it is the VTEP's only when both share a Layer 2 segment. Each router on the way then writes a new outer Ethernet header, as for any routed packet.
  • Across a VXLAN segment, does VM-A's traceroute to VM-B show the underlay's leaf and spine hops?
    No. Within one segment the VTEPs bridge the frame, and RFC 8014 says the tenant packet's TTL or Hop Limit is not modified; only the outer header's TTL is decremented by underlay routers. The tenant's probes therefore never expire in the underlay, and VM-B answers as if it were on the same LAN, while the leaves and spine stay invisible.

saying these in an interview costs you the question

  • The source VTEP rewrites the inner destination MAC to the remote VTEP's MAC address.
  • Spine switches must learn the tenant VMs' MAC addresses to forward VXLAN traffic.
  • The outer destination MAC is always the remote VTEP's own MAC address.
  • The remote VTEP routes the inner packet and decrements its TTL before delivery.
  • The original frame's FCS travels inside the tunnel and is checked by the receiving VM.