How do underlay IP multicast and ingress replication differ as ways for a VXLAN VTEP to deliver BUM traffic, and what does each cost?
answer
- who makes the copies
- network state versus edge bandwidth
- one group per VNI
- N minus one copies
basics
~20 sWith underlay multicast, a VTEP sends one copy to the VNI's group and the routed underlay replicates it, which needs multicast routing state. With ingress replication, the VTEP sends one unicast copy per peer VTEP, trading underlay state for source bandwidth.
solid answer
~40 sRFC 7348's model maps each VNI to an IP multicast group through the management layer. VTEPs join it with IGMP, using `(*,G)` joins because the set of senders keeps changing, and the underlay — a multicast routing protocol such as PIM Sparse Mode, or bidirectional PIM, which the RFC calls more efficient since every VTEP both sends and receives — builds the tree. The source sends one packet. **Ingress (head-end) replication** needs no multicast in the underlay: the source VTEP keeps a list of peer VTEPs for the VNI and sends N−1 unicast copies, so a VNI spanning 24 VTEPs costs 23 copies of every flood on the source's uplinks. RFC 7365 calls this a **bandwidth versus state** trade-off. Without a control plane, the peer list is static configuration on every VTEP.
go deeper
Recall the two delivery methods for VXLAN floods and who makes the copies in each: the underlay's routers, or the sending VTEP.
Explain group mapping, IGMP joins and the multicast tree, then compute N minus one copies for ingress replication and say what each demands from the underlay.
Show the operational side: multicast trees and group mappings to debug, versus peer lists to keep consistent on every VTEP, and how flood volume multiplies at the edge.
Weigh underlay multicast state against edge bandwidth for the fabric's size and BUM volume, and treat both as stopgaps for a design whose real cost is the flooding.
## Why a flood needs replication at all A **VTEP** (VXLAN Tunnel End Point) must send each **BUM** frame — broadcast, unknown unicast and multicast — to every other VTEP that hosts the same **VNI** (VXLAN Network Identifier). The underlay is a routed IP network that only knows how to deliver packets to IP addresses, so *someone* has to turn one frame into many packets. Either the **network** does it (IP multicast) or the **sending VTEP** does it (ingress replication). RFC 7365, the IETF framework for data-centre network virtualization, lists exactly these two methods for BUM traffic. ## Option 1: underlay IP multicast RFC 7348 §4.2 describes this as the flood method of its data-plane learning scheme: - **One group per VNI.** The management layer maps each VNI to an IP multicast group and hands the mapping to every VTEP over a management channel. - **VTEPs join.** A VTEP that hosts the VNI sends IGMP membership reports to its upstream router, so the underlay knows it wants the group. - **Shared trees.** VTEPs use `(*,G)` joins — any source, one group — because the set of VTEPs sending into the segment is unknown and changes as hosts come and go. - **A multicast routing protocol** such as PIM Sparse Mode builds the tree. RFC 7348 adds that bidirectional PIM would be more efficient, because every VTEP is both a source and a receiver for its groups. - **Layer 2 underlays work too:** RFC 7348 §4.3 notes VXLAN can run over a layer 2 network, where IGMP snooping keeps the replication efficient. The source VTEP sends **one** encapsulated packet, addressed to the group; routers copy it where the tree branches. ## Option 2: ingress replication Here the source VTEP holds a **peer list** for each VNI and sends a separate **unicast** VXLAN packet to every VTEP on it. The underlay sees only unicast traffic, so it needs no multicast routing, no rendezvous point and no group state. In a flood-and-learn design, nothing discovers the list — it is configured on every VTEP, and a VTEP missing from one list silently misses that VTEP's floods. ## The arithmetic A VNI spans **24 VTEPs**, and one VTEP sources **2 Mb/s** of BUM traffic in it (ARP, unknown unicast, tenant multicast), before encapsulation overhead: 1. **Multicast:** 1 copy leaves the source → 2 Mb/s on its uplinks; the underlay forwards one copy per tree branch. 2. **Ingress replication:** 24 − 1 = **23 copies** → 23 × 2 = **46 Mb/s** on the source's uplinks, and the same 23 copies cross the fabric as separate unicast packets. The ingress cost grows linearly with the number of VTEPs in the VNI, and every VTEP that floods pays it. ## Side by side | | Underlay multicast | Ingress replication | |---|---|---| | Who copies | Underlay routers, along a tree | The source VTEP | | Packets leaving the source | 1 per flood | N − 1 per flood | | Underlay requirement | Multicast routing and per-group state | Unicast routing only | | Membership | VTEPs join the VNI's group | Peer list configured per VNI | | Suits | Many VTEPs per VNI, heavy BUM | Few VTEPs, light BUM | | Typical failure | Tree or group mapping broken | A peer missing from a list | RFC 7365 §4.2.3 states the rule of thumb: there is a bandwidth versus state trade-off; when the number of hosts per group is large, underlay multicast trees may be more appropriate, and when it is small (it gives 2–3 as the example) or the multicast traffic is light, ingress replication may not be an issue. ## Sharing groups between VNIs A 24-bit VNI allows 16,777,216 segments; far more than an underlay's multicast state is normally sized for. RFC 7348 leaves the VNI-to-group mapping to management, so operators can put several VNIs on one group. That cuts underlay state, but every VTEP that joins the group receives floods for **all** the VNIs on it and must discard those for segments it does not host — RFC 7365 lists shared rather than dedicated underlay multicast trees as a possible trade-off in the same spirit. Either way, the flood still reaches every VTEP in the segment, which is the scaling problem replication methods cannot remove.
- Why does RFC 7348 use (*,G) joins, and why does it mention bidirectional PIM?The set of VTEPs sending into a VNI is unknown and changes as hosts start and stop, so a VTEP joins the group for any source rather than per source. Because every VTEP is both a sender and a receiver on its groups, RFC 7348 notes that bidirectional PIM, which builds one shared bidirectional tree, would be more efficient.
- What happens if two VTEPs map the same VNI to different multicast groups?Each floods to a group the other never joined, so BUM frames between them, ARP requests included, never arrive. Hosts behind them cannot resolve each other, although the VTEPs reach each other fine over unicast. The mapping comes from the management layer, so the fix belongs there.
- When is ingress replication the better choice despite its bandwidth cost?When few VTEPs share each VNI or BUM traffic is light, the N−1 copies are cheap, and the underlay avoids multicast routing, rendezvous points and per-group state entirely. RFC 7365 makes the same point, citing groups of two or three hosts.
saying these in an interview costs you the question
- Ingress replication needs multicast routing in the underlay, just like multicast mode.
- With ingress replication the spine routers make the copies.
- In multicast mode the source sends one unicast copy to each VTEP.
- RFC 7348 requires a dedicated multicast group for every VNI.
- Ingress replication costs the same however many VTEPs host the VNI.