skip to content

VXLAN

VXLAN wraps Ethernet frames in UDP so a Layer 2 segment can span a routed fabric, with a 24-bit VNI naming each segment. Interviewers use it for data-centre and container networking.

on this pageshow

explore

questions

page 1 of 2

Why does VXLAN identify segments with a 24-bit VNI rather than VLAN IDs, and how many segments does that allow?

level: juniorimportance: must knowfreq 42%

answer

  1. too many tenants for one tag space
  2. count the identifier bits
  3. each identifier is its own broadcast domain
  4. the VTEP adds it, the host never sees it

basics

~20 s

VXLAN carries a 24-bit VXLAN Network Identifier in its own header, giving 2^24 = 16,777,216 segment IDs against 4,094 usable 12-bit VLAN IDs, so a shared data centre can give every tenant many isolated layer 2 segments.

solid answer

~40 s

An 802.1Q VLAN ID is 12 bits, and with 0 and 4095 reserved by the IEEE only 4,094 VLANs remain for the whole switched domain. RFC 7348 calls that limit inadequate once a provider hosts many tenants, each wanting several segments and each choosing VLAN IDs and MAC addresses independently. VXLAN instead puts a 24-bit **VXLAN Network Identifier (VNI)** in its own header, so one administrative domain can hold up to 2^24 = 16,777,216 segments ("up to 16 M" in the RFC). Each VNI is a separate layer 2 broadcast domain: a frame is delivered only to hosts on the same VNI, so overlapping MAC addresses in different segments never cross over. Hosts never see the VNI; the VTEP adds it on encapsulation and removes it on decapsulation.

go deeper

for a junior

Recall the two widths and their counts: 12-bit VLAN ID with 4,094 usable values, 24-bit VNI with 16,777,216. Say that each VNI is its own broadcast domain.

for a middle

Explain who stamps the VNI and where: the ingress VTEP writes it into the VXLAN header from the local VLAN or port, and the egress VTEP delivers only within that VNI.

for a senior

Point out that the identifier space is rarely the real limit: per-switch VNI, MAC and ARP table sizes bind first, and a routed tenant consumes extra VNIs.

for a principal

Frame the VNI as one layer of tenancy: bridging isolation by VNI, routing isolation by per-tenant VRF, and the allocation plan that keeps both consistent across a fabric.

## The problem VLAN IDs could not solve A **VLAN** splits one switched Ethernet network into separate **broadcast domains**. Each frame on a trunk carries an IEEE 802.1Q tag whose **VLAN ID** field is 12 bits wide. Twelve bits give 4,096 values; the IEEE reserves 0 and 4095, which leaves **4,094 usable VLAN IDs** (the 802.1Q values are the IEEE's, not an RFC's). RFC 7348, the Informational RFC that defines VXLAN, lists why that ceiling hurt multi-tenant data centres: - **Tenant count.** A provider serving many customers, each needing several segments, runs out of 4,094 IDs quickly; the RFC says the limit is "often inadequate". - **Independent numbering.** Tenants assign their own MAC addresses and VLAN IDs, so the same values collide on the shared physical network. - **Spanning tree.** Large layer 2 domains depend on spanning tree, which blocks redundant links, and the RFC notes several data centres limit how many VLANs they use because of it. ## What the VNI is VXLAN is a **layer 2 overlay on a layer 3 network**: an Ethernet frame is wrapped in an outer IP/UDP packet and carried across a routed fabric between two **VXLAN Tunnel End Points (VTEPs)**. The wrapper includes an 8-byte VXLAN header, and the identifier that matters for tenancy sits in it: the 24-bit **VXLAN Network Identifier (VNI)**, also called the VXLAN segment ID. Each value names one **VXLAN segment**, an independent layer 2 broadcast domain. ## The arithmetic | Identifier | Field width | Values | Usable for segments | |---|---|---|---| | 802.1Q VLAN ID | 12 bits | 4,096 | 4,094 (0 and 4095 reserved by the IEEE) | | VXLAN VNI | 24 bits | 16,777,216 | "up to 16 M" per RFC 7348 | 2^24 = 16,777,216, which is 4,096 times the 12-bit space. RFC 7348 phrases the result as up to 16 million segments coexisting within the same **administrative domain**. ## How the VNI isolates segments 1. A host sends an ordinary Ethernet frame; it knows nothing about VXLAN. 2. The ingress VTEP decides which segment the frame belongs to, usually from the local VLAN or port it arrived on, and writes that segment's VNI into the header. 3. The egress VTEP reads the VNI and delivers the inner frame only to local hosts attached to the same VNI. Because every MAC lookup is made inside one VNI, RFC 7348 notes you "could have overlapping MAC addresses across segments but never have traffic cross over". Two tenants can therefore reuse the same MAC addresses, and the same VLAN numbers on their own ports, without colliding. ## What the VNI does not do on its own - **It does not make every tenant routable.** A VNI is a bridging domain. Routing between a tenant's subnets, and keeping overlapping IP prefixes apart, needs a per-tenant routing instance (a VRF), usually tied to its own VNI. - **It does not remove hardware limits.** MAC tables, ARP tables and the number of VNIs a single switch can hold are implementation limits, and they usually bind long before 16 million. - **It is not 16 million tenants.** A tenant normally uses several segments, plus one more VNI for its routing instance in a routed design. - **It does not replace VLANs at the edge.** Servers still send untagged or VLAN-tagged frames to the leaf switch; the VLAN simply becomes a local attachment detail that the VTEP maps to a VNI. ## Where it sits among overlays VXLAN is not the only overlay with a wider identifier: NVGRE (RFC 7637) carries a 24-bit Virtual Subnet Identifier in a GRE key, and Geneve (RFC 8926) carries a 24-bit VNI as well. The interview point is the same for all three: a 24-bit tenant segment identifier carried in an encapsulation header, outside the tenant's own frame, lifts the 4,094-segment ceiling that a 12-bit tag imposed.

  • Does a VXLAN fabric stop using VLANs altogether?
    No. Servers and appliances still send untagged or 802.1Q-tagged frames to the leaf. The VTEP maps the local VLAN, or the port, to a VNI and, as RFC 7348 section 6.1 recommends, strips the tag before encapsulating. The VLAN becomes a local attachment detail on one switch; the VNI is the segment's identity across the fabric.
  • Why are 16 million VNIs not the same as 16 million tenants?
    A tenant usually needs several bridged segments, and in a routed design one more VNI for its routing instance. In practice the binding limit is hardware: how many VNIs, MAC entries and ARP entries one switch can hold. The 24-bit space removes the identifier ceiling, not the resource ceiling.

saying these in an interview costs you the question

  • VXLAN widens the 802.1Q VLAN ID to 24 bits inside the same tag.
  • A 24-bit VNI means a fabric can host 16 million tenants without other limits.
  • Two tenants cannot reuse the same MAC addresses inside one VXLAN fabric.
  • Each host must be configured with the VNI of its segment.
  • VXLAN segments all share one broadcast domain separated only by IP subnets.
open as a page

What is a VXLAN Tunnel Endpoint (VTEP), and what does it do to an Ethernet frame entering and leaving the tunnel?

level: juniorimportance: must knowfreq 42%

basics

~20 s

A VTEP originates and terminates VXLAN tunnels: it wraps a host's Ethernet frame in a VXLAN header carrying the segment's VNI, plus UDP and an outer IP header addressed VTEP to VTEP, and the far VTEP strips them and delivers the original frame.

open as a page

How many bytes does VXLAN encapsulation add to an Ethernet frame over IPv4, and how does that change with IPv6 or an outer 802.1Q tag?

level: middleimportance: must knowfreq 36%

basics

~20 s

Over IPv4, VXLAN adds 50 bytes: outer Ethernet 14, outer IPv4 20, UDP 8 and the VXLAN header 8. An outer 802.1Q tag makes it 54; an outer IPv6 header, 40 bytes, makes it 70, or 74 tagged.

open as a page

Why do VXLAN fabrics adopt BGP EVPN as a control plane, and what does it advertise that flood-and-learn must discover by flooding?

level: middleimportance: must knowfreq 30%

basics

~20 s

BGP EVPN lets each VTEP announce the MAC and IP addresses it learned locally as BGP routes, so remote VTEPs know where hosts live before traffic flows, instead of learning them by flooding unknown and ARP traffic across the fabric.

open as a page

In a VXLAN segment using data-plane learning, what does each VTEP learn as two hosts complete their first ARP exchange?

level: middleimportance: must knowfreq 32%

basics

~20 s

Host A's ARP broadcast is flooded to every VTEP in the VNI, so each learns A's MAC behind A's VTEP. B's unicast reply goes to A's VTEP alone, which learns B's MAC behind B's VTEP; afterwards both directions are known unicast.

open as a page

When a VXLAN overlay over IPv4 gives tenant hosts a 1,500-byte MTU, what IP MTU must every underlay link carry, and why?

level: middleimportance: must knowfreq 38%

basics

~20 s

At least 1,550 bytes: the 1,500-byte tenant packet plus its 14-byte inner Ethernet header, 8 bytes of VXLAN, 8 of UDP and 20 of outer IPv4. RFC 7348 forbids VTEPs to fragment, so a smaller link drops large packets.

open as a page

In VXLAN, how do two VTEPs carry a frame from a VM on rack 1 to a VM on rack 2 in the same segment?

level: middleimportance: must knowfreq 32%

basics

~20 s

The source VTEP maps the destination MAC in the VM's VNI to a remote VTEP and wraps the frame in VXLAN, UDP and an outer VTEP-to-VTEP IP header; the underlay routes that packet, and the remote VTEP unwraps it and delivers the unchanged frame.

open as a page

In BGP EVPN for VXLAN, what do route types 2, 3 and 5 each carry, and when does a VTEP originate each one?

level: seniorimportance: must knowfreq 20%

basics

~20 s

Type 2 (MAC/IP Advertisement) carries a host's MAC and optional IP behind a VTEP; type 3 (Inclusive Multicast Ethernet Tag) enrols a VTEP for a VNI's BUM traffic; type 5 (IP Prefix, RFC 9136) carries a prefix without a MAC.

open as a page

Why does VXLAN flood-and-learn scale poorly in a large data-centre fabric, and which costs grow as VTEPs, VNIs and hosts are added?

level: seniorimportance: must knowfreq 24%

basics

~20 s

Every VTEP in a VNI carries every BUM frame; ARP and unknown unicast flood because nothing knows mappings in advance; the underlay holds multicast state or sources send N−1 copies; and a moved host is relearned only when it sends.

open as a page

After a VXLAN overlay goes live, logins and small requests between VMs work but large transfers stall; why doesn't Path MTU Discovery rescue the VMs, and what fixes it?

level: seniorimportance: must knowfreq 30%

basics

~20 s

Full-size packets no longer fit once VXLAN adds 50 bytes, and any ICMP error goes to the VTEP, the outer source, never the VM, so the VM's Path MTU Discovery never learns. Fix the underlay MTU, or lower the tenant MTU.

open as a page

Two tenants in one VXLAN fabric both use 10.0.0.0/24; how does the fabric keep their bridging and routing apart?

level: seniorimportance: must knowfreq 28%

basics

~20 s

Every lookup happens inside a tenant context: frames ride in the tenant's own L2 VNI, routed packets are looked up in the tenant's own VRF and carried in its L3 VNI, so tenant A's 10.0.0.5 and tenant B's are never compared.

open as a page

In VXLAN, which headers wrap the original Ethernet frame, from the outside in, and what job does each one do?

level: juniorimportance: should knowfreq 28%

basics

~20 s

VXLAN wraps the original Ethernet frame, minus its FCS, in an 8-byte VXLAN header carrying the 24-bit VNI, then UDP to port 4789, an outer IP header between the two VTEPs, and an outer Ethernet header with a new FCS.

open as a page

In a VXLAN overlay without a control plane, what does flood-and-learn mean, and what is BUM traffic?

level: juniorimportance: should knowfreq 22%

basics

~20 s

Flood-and-learn is VXLAN's data-plane learning: a VTEP floods any frame it cannot send to one known VTEP to every VTEP in the segment, and learns MAC-to-VTEP mappings from frames it decapsulates. BUM means broadcast, unknown-unicast and multicast frames.

open as a page

How does a VXLAN VTEP set the outer UDP header's destination port, source port and checksum under RFC 7348?

level: middleimportance: should knowfreq 16%

basics

~20 s

Destination port 4789, IANA's assignment, is the default but should be configurable; the source port is the VTEP's choice, recommended as a hash of inner-frame fields in 49152-65535; the checksum SHOULD be zero, and receivers MUST accept zero.

open as a page

What does each field of the 8-byte VXLAN header carry, and what must a VTEP do with its flag and reserved bits?

level: middleimportance: should knowfreq 18%

basics

~20 s

The VXLAN header is 8 bytes: 8 flag bits whose I bit must be 1 for a valid VNI, 24 reserved bits, the 24-bit VNI, then 8 more reserved bits. Reserved bits are sent as zero and ignored on receipt.

open as a page

How do underlay IP multicast and ingress replication differ as ways for a VXLAN VTEP to deliver BUM traffic, and what does each cost?

level: middleimportance: should knowfreq 27%

basics

~20 s

With underlay multicast, a VTEP sends one copy to the VNI's group and the routed underlay replicates it, which needs multicast routing state. With ingress replication, the VTEP sends one unicast copy per peer VTEP, trading underlay state for source bandwidth.

open as a page

For a VXLAN overlay, what do you trade when you raise the underlay MTU instead of lowering every tenant's MTU to 1,450 bytes?

level: middleimportance: should knowfreq 26%

basics

~20 s

Raising the underlay MTU leaves tenants at the standard 1,500 bytes but needs every underlay link on every path changed. Lowering tenants to 1,450 fits any 1,500-byte underlay but needs every tenant host changed, and a missed host fails silently.

open as a page

In a VXLAN fabric that routes for its tenants, what is the difference between an L2 VNI and an L3 VNI?

level: middleimportance: should knowfreq 24%

basics

~20 s

An L2 VNI names one bridged subnet, so the receiving VTEP looks up the destination MAC in that bridge table; an L3 VNI names a tenant's VRF, so the receiving VTEP looks up the destination IP in that tenant's routing table.

open as a page

On a VXLAN VTEP, how is a local VLAN mapped to a VNI, and why may two VTEPs use different VLAN IDs for one segment?

level: middleimportance: should knowfreq 30%

basics

~20 s

The VTEP holds a configured VLAN-to-VNI table; it strips the 802.1Q tag before encapsulating and maps the arriving VNI back to its own local VLAN, so with the usual domain-wide VNIs only the VNI must agree and VLAN IDs stay locally significant.

open as a page

Where can a VXLAN VTEP run, in a hypervisor's virtual switch or in a top-of-rack switch, and what does each placement trade?

level: middleimportance: should knowfreq 22%

basics

~20 s

A software VTEP in the hypervisor's virtual switch knows each VM's attachment directly but spends server CPU; a hardware VTEP in a top-of-rack switch encapsulates at line rate and also serves bare-metal and VLAN-only hosts, within the switch's table limits.

open as a page

A capture shows VXLAN packets reaching the remote VTEP, yet the tenant hosts never see the frames; which encapsulation-format mismatches would you check?

level: seniorimportance: should knowfreq 12%

basics

~20 s

Walk the packet against RFC 7348: a destination port the receiver does not decapsulate, a clear I bit or unknown VNI, an inner VLAN tag the receiver discards by default, or a UDP checksum it rejects.

open as a page

In a BGP EVPN VXLAN fabric, why does every leaf carry the same gateway IP and MAC, and how does EVPN track a host that moves between leaves?

level: seniorimportance: should knowfreq 18%

basics

~20 s

A distributed anycast gateway puts one gateway IP and MAC on every leaf, so hosts are routed at their own leaf and keep a valid gateway ARP entry after moving; EVPN tracks moves with MAC Mobility sequence numbers on type-2 routes.

open as a page

How does ARP suppression on a BGP EVPN VXLAN VTEP answer ARP requests locally, and which requests still get flooded across the fabric?

level: seniorimportance: should knowfreq 16%

basics

~20 s

The VTEP keeps a proxy table of IP-to-MAC bindings, learned by snooping local ARP and from remote type-2 MAC/IP routes, and answers a local ARP request itself on a hit; a request it cannot answer is still flooded to remote VTEPs.

open as a page

How does BGP EVPN multihome a server to two VXLAN leaves through an Ethernet segment, and what do route types 1 and 4 do?

level: seniorimportance: should knowfreq 12%

basics

~20 s

Both leaves give the server's link bundle one Ethernet Segment Identifier; type-4 Ethernet Segment routes let them discover each other and elect a designated forwarder for BUM, and type-1 Ethernet A-D routes provide aliasing and mass withdrawal.

open as a page

Why does a VXLAN VTEP derive the outer UDP source port from the inner packet's headers, and what happens in a leaf-spine underlay when every flow gets one port?

level: seniorimportance: should knowfreq 22%

basics

~20 s

Between two VTEPs every outer field except the UDP source port is fixed, so a per-flow hash carried in that port is the underlay ECMP's only entropy. With one port for every flow, each VTEP pair's traffic rides a single path.

open as a page

Under RFC 9135, how do symmetric and asymmetric integrated routing and bridging differ in a VXLAN fabric, and what does each demand of every leaf?

level: seniorimportance: should knowfreq 15%

basics

~20 s

Asymmetric IRB routes only at the ingress VTEP and bridges into the destination subnet's L2 VNI, so every leaf carries every subnet; symmetric IRB routes at both VTEPs over the tenant's L3 VNI, so each leaf carries only its local subnets.

open as a page

What can an underlay router see of VXLAN traffic between two VTEPs, and how does that shape underlay filtering and troubleshooting?

level: seniorimportance: should knowfreq 18%

basics

~20 s

An underlay router forwards on the outer headers: VTEP source and destination addresses, UDP port 4789 and a hashed source port. Tenant addresses sit inside the payload, so outer-header filters act per VTEP pair, and tenant-level debugging starts at the VTEPs.

open as a page

Designing the IP underlay for a multi-tenant VXLAN fabric of 40 leaves, what must it give the VTEPs, and which routing, MTU and multicast choices would you make?

level: principalimportance: should knowfreq 14%

basics

~10 s

A routed leaf-spine with no spanning tree; a routing protocol carrying only VTEP addresses and links, with full-width ECMP and no summarisation; one fabric-wide jumbo MTU; and unicast-only BUM handling unless flood-and-learn needs multicast.

open as a page

After an underlay change in a flood-and-learn VXLAN fabric, hosts in one VNI resolve ARP only within a rack while established flows keep working; what is broken?

level: seniorimportance: nice to knowfreq 12%

basics

~20 s

The BUM path is broken, not the tunnel: learned unicast still flows between VTEPs, but floods no longer cross racks — typically underlay multicast for the VNI's group, a VNI-to-group mismatch, or a missing ingress-replication peer.

open as a page

Why does a hardware VXLAN VTEP usually source its tunnels from a loopback address, and what breaks when that address is missing from underlay routing?

level: seniorimportance: nice to knowfreq 12%

basics

~20 s

A loopback stays reachable while any uplink survives, so it makes a stable VTEP address. It is every tunnel's outer source and destination; if the underlay lacks its route, return traffic to that VTEP is dropped while its outbound traffic still flows.

open as a page

showing 1–30 of 31