Where can a VXLAN VTEP run, in a hypervisor's virtual switch or in a top-of-rack switch, and what does each placement trade?
answer
- same role, two boxes
- server CPU versus switch tables
- bare-metal hosts need a switch
- offloads the new header defeats
- gateway between VLAN and VNI
basics
~20 sA software VTEP in the hypervisor's virtual switch knows each VM's attachment directly but spends server CPU; a hardware VTEP in a top-of-rack switch encapsulates at line rate and also serves bare-metal and VLAN-only hosts, within the switch's table limits.
solid answer
~50 sRFC 7348 places the VTEP in the hypervisor in its examples but allows it on a physical switch or server, in software or hardware. A **software VTEP** in the hypervisor's virtual switch learns directly when a VM attaches or detaches, needs nothing special from the physical network, and makes every server a tunnel endpoint; RFC 8014 notes the price: server CPU for every packet, and a new header that can disable network-adapter offloads unless encapsulation is offloaded too. A **hardware VTEP** in the top-of-rack switch encapsulates in the switch's forwarding hardware, so servers send plain or VLAN-tagged frames and bare-metal servers, appliances and legacy hosts can join a segment; the costs are fixed hardware table sizes and the need to tell the switch which access-port VLAN maps to which segment. A switch acting as a **VXLAN gateway** is the hardware case that bridges a VLAN into a VNI.
go deeper
Recall the two homes of a VTEP, the hypervisor's virtual switch and the top-of-rack switch, and that the role is identical in both.
Explain what each placement trades: attachment knowledge and CPU on the server side, line rate, bare-metal reach and table limits on the switch side, plus the gateway's three rules.
Decide placement from the estate: share of bare-metal hosts, server CPU headroom, adapter offload support and switch table sizes, and plan how the two kinds coexist.
Weigh who owns the overlay boundary: hypervisor VTEPs give it to the compute platform, switch VTEPs to the network team, and the choice shapes change control and fault ownership.
## The same role in two places A **VTEP** does one job wherever it runs: map a host's frame to a VXLAN segment, encapsulate it toward a remote VTEP address, and decapsulate what arrives. **RFC 7348** draws its examples with the VTEP inside the **hypervisor** of the server hosting the virtual machines, then says plainly that a VTEP could also be on a physical switch or physical server and could be implemented in software or hardware. **RFC 8014**, the IETF's overlay architecture, discusses the same choice under the name **Network Virtualization Edge (NVE)**. Where the VTEP runs decides where the overlay begins, who has to know about it, and which resources pay for it. ## Software VTEP in the hypervisor Here the hypervisor's virtual switch does the encapsulation, so the overlay begins inside the server. - **Direct knowledge of attachment.** The hypervisor knows the moment a VM is created, moved or removed; RFC 8014 notes that with the NVE inside the hypervisor, no on-the-wire protocol between host and NVE needs standardising. - **No demands on the physical switches.** The physical network carries the overlay as server-to-server UDP/IP traffic between VTEP addresses. - **Scale follows the servers.** Every server is a VTEP, so the number of tunnel endpoints equals the number of virtualised servers, which enlarges the set of VTEPs that flooded traffic must reach. - **CPU cost.** RFC 8014 states that implementing the NVE entirely on a server spends server CPU, and that the added overlay header can disable existing network-adapter offloads that are not prepared for it. RFC 8014 suggests offloading encapsulation and decapsulation onto the adapter to win that back. - **No help for hosts without the hypervisor.** A bare-metal server or a physical appliance has no hypervisor VTEP, unless it runs a VTEP itself. ## Hardware VTEP in the top-of-rack switch Here the server sends ordinary frames and the switch it plugs into encapsulates them. - **Line-rate encapsulation** in the switch's forwarding hardware, with no server CPU spent on it. - **Any host can join.** Bare-metal servers, storage, physical firewalls and legacy machines become members of a segment just by being on the right access port or VLAN. - **The access link carries tenant frames.** RFC 8014's **split-NVE** case, where a hypervisor hands encapsulation to the adjacent switch, describes that link: traffic of a given tenant system is tagged with a VLAN C-TAG on the access link, and that tag identifies which virtual network it joins, so host and switch must agree on the tag for each segment. - **Fewer endpoints.** One VTEP per rack instead of one per server keeps the number of tunnel endpoints small. - **Hardware limits.** Remote MAC-to-VTEP entries and VNI mappings live in fixed-size forwarding tables; how large they are is an implementation property of the switch, and it bounds how many hosts and segments the rack can serve. ## The gateway VTEP RFC 7348 describes a switch acting as a **VXLAN gateway** that connects a VXLAN segment to hosts on an ordinary VLAN: 1. A frame arriving **from the overlay** is decapsulated and forwarded to a physical port based on its inner destination MAC. 2. A decapsulated frame carrying an **inner VLAN tag** SHOULD be discarded unless the gateway is configured to pass it. 3. A frame arriving **from a VLAN port** is mapped to a VXLAN segment by its VLAN ID, and that VLAN ID is removed before encapsulation unless configured otherwise. RFC 7348 adds that gateways may be top-of-rack, core or WAN-edge devices, in software or hardware. ## Side by side | Concern | Software VTEP (hypervisor) | Hardware VTEP (top-of-rack) | |---|---|---| | Where the overlay begins | inside the server | at the leaf switch | | Who pays per packet | server CPU, unless offloaded | switch forwarding hardware | | Learns VM attachment from | the hypervisor directly | the access port and VLAN tag | | Bare-metal and appliances | not served | served | | Number of VTEPs | one per server | one per rack | | Main limit | CPU and lost adapter offloads | hardware table sizes | ## Mixing both The underlay does not care which kind of VTEP sent a packet: both produce the same UDP/IP encapsulation between VTEP addresses. A fabric can therefore mix them, with hypervisor VTEPs for virtual machines and switch VTEPs or gateways for physical hosts, as long as every VTEP agrees on segment membership and on the mapping of MAC addresses to VTEP addresses.
- How does a hardware VXLAN gateway let a bare-metal server on a VLAN talk to virtual machines on a VXLAN segment?Per RFC 7348, the gateway maps frames arriving on the server's VLAN to a VXLAN segment by their VLAN ID, removes the tag and encapsulates them toward the VMs' VTEPs. In the other direction it decapsulates and forwards on the inner destination MAC to the physical port, discarding frames with an inner VLAN tag unless configured otherwise. The server needs no VXLAN support.
- Why can adding a VXLAN header slow a software VTEP even on a fast server?RFC 8014 points to two costs: encapsulation and decapsulation run on server CPU for every packet, and the extra overlay header can disable network-adapter offloads, such as checksum and TCP offloads, that are not prepared for it. RFC 8014 suggests offloading encapsulation and decapsulation onto the adapter to win that performance back.
saying these in an interview costs you the question
- A VXLAN VTEP must be a hardware switch, because hypervisors cannot encapsulate.
- With top-of-rack VTEPs, servers must still put VXLAN headers on their own frames.
- Hardware VTEPs need no table mapping remote MAC addresses to VTEP addresses.
- A gateway should pass decapsulated frames that carry an inner VLAN tag by default.
- Software VTEPs cost nothing beyond the few header bytes they add.