In SD-WAN, what is the difference between the overlay and the underlay, and why does that split let a branch use almost any transport?
answer
- two layers, two jobs
- rented circuits versus built tunnels
- private prefixes live in one layer only
- underlay needs only edge-to-edge reachability
- quality and shared fate are inherited
basics
~20 sThe underlay is the rented transports, such as MPLS, broadband or LTE, that only carry packets between edge addresses; the overlay is the encrypted tunnels, routes and policy built over them, so sites see one private network whatever the transport.
solid answer
~50 sThe **underlay** is whatever a site rents: an MPLS VPN, a broadband or dedicated internet access (DIA) circuit, an LTE or 5G modem. Each one gives the SD-WAN edge a transport address and reachability to other edges, and nothing more. The **overlay** is what the edges build on top: encrypted tunnels across every usable transport, the company's private prefixes exchanged through central controllers and pointed into those tunnels, and the policy deciding which application uses which tunnel. Because private routing exists only in the overlay, any transport that offers IP reachability becomes usable, and a circuit can be added or swapped without touching the private routing plan. The catch is that the overlay inherits each underlay's loss, latency and shared fate, so the edges must measure every path and the design must keep the underlays physically independent.
go deeper
Recall the two layers: the underlay is the rented circuits, the overlay is the tunnels, routes and policy the edges build over them. Be able to name the usual transports.
Explain why private prefixes living only in the overlay makes any IP transport usable, and why tunnels can only pair transports that reach each other.
Show that the overlay inherits loss and shared fate from the underlay, and that physical diversity of circuits matters more than how many tunnels exist.
Frame the split as a trade: circuits become commodity pipes, while privacy, routing and quality judgement move into edge software and a controller you must now run.
## Two layers, two jobs An **SD-WAN** (software-defined wide-area network) connects branch sites, such as stores, offices or clinics, to each other and to data centres by splitting the WAN into two layers that are bought, designed and operated separately. No IETF standard defines SD-WAN: it is an architecture that many implementations share, so the terms below describe that architecture, not a product. - The **underlay** is the set of transports a site actually rents: an MPLS VPN from a carrier, a consumer **broadband** line, a **dedicated internet access (DIA)** circuit, an **LTE or 5G** modem. Each gives the site's edge device an IP address, its **transport address**, and some way to reach other edges. - The **overlay** is what the SD-WAN edge devices build on top of those transports: encrypted tunnels between edges, the company's private routes carried through those tunnels, and the policy that decides which application uses which tunnel. ## What the underlay provides | Transport | What the site gets | Typical character | |---|---|---| | MPLS VPN | Private reachability inside one carrier's network | Contracted loss and latency targets, costly per Mb/s, slow to order | | Broadband | Internet reachability | Cheap, high bandwidth, often asymmetric, best effort, shared | | DIA | Internet reachability on a dedicated circuit | Symmetric, with a contract that ends at the ISP's network | | LTE / 5G | Internet reachability over radio | Quick to install, often metered, variable latency | In an SD-WAN the underlay's job shrinks to one thing: carry packets between edges' transport addresses. It does not need to know the company's private prefixes, its addressing plan or its applications. ## What the overlay adds 1. **Tunnels.** Each edge builds encrypted tunnels to the edges it must reach, one per usable pair of transports. The encapsulation is commonly IPsec or something similar; how those tunnels are framed and keyed is a separate subject. 2. **Private routing.** Edges learn each other's site prefixes (a store's `10.20.31.0/24`, from RFC 1918 private space) over a control channel to central controllers, and install them pointing into the tunnels. 3. **Policy.** Which applications may use which transport, and in what order of preference. 4. **Measurement.** Edges probe every tunnel for loss, latency and jitter, because the overlay can only be as good as the path under it. ## Why the split matters Because private routing lives only in the overlay: - **Any transport that offers IP reachability becomes usable.** A store can open on an LTE modem on day one and add broadband when the line is installed, without changing its private routing. - **Transports can be mixed and replaced.** Swapping one ISP for another changes a tunnel's endpoint address, not the store's private prefix. - **The carrier stops carrying company routes.** In a pure MPLS VPN the provider's network holds the customer's routes; with an overlay the provider sees only encrypted traffic between edge addresses. - **Policy is defined once and pushed everywhere.** The same per-application rules reach three hundred stores from a central point instead of being typed into each router. ## Pairing tunnels across transports A tunnel can only form where the two underlays reach each other. An MPLS VPN is a private routing domain, so an address on a hub's MPLS port is normally not reachable from the internet. A store with only broadband therefore builds its tunnel to the hub's internet-facing transport, not to the hub's MPLS port, while a store with both transports builds one tunnel over each. Implementations group transports into named classes so that edges only attempt pairs that can work. This is also why hubs and data centres usually keep both an MPLS and an internet transport: they must be reachable from stores that have only one of them. ## What the overlay cannot change - **It inherits the underlay's quality.** A broadband line dropping 2 % of packets drops them inside the tunnel as well. The overlay can steer an application onto a better path; it cannot repair the bad one. - **It inherits shared fate.** If the MPLS circuit and the broadband line enter the building through the same duct, one cut removes both underlays and every tunnel riding on them. - **It adds overhead.** Every tunnel adds headers, so less payload fits in the path MTU; handling that with fragmentation avoidance or MSS clamping is a separate subject. - **It adds a control dependency.** The overlay's routes and keys come from controllers, so their reachability and redundancy become part of the design. The split is the whole idea of SD-WAN: treat circuits as interchangeable pipes, put intelligence and privacy in software at the edges, and accept that the result is only as reliable as the independence and quality of the pipes underneath.
- Why can a store's broadband-only SD-WAN edge not build a tunnel to the hub's MPLS port?An MPLS VPN is a private routing domain inside one carrier, so the hub's MPLS address is not normally reachable from the internet. Tunnels form only between transports that can reach each other, so the broadband-only store tunnels to the hub's internet-facing transport. That is why hubs keep both kinds of transport.
- With 300 stores, why do designs rarely build a full mesh of tunnels on every transport?A full mesh needs n(n-1)/2 tunnels per transport: 300 x 299 / 2 = 44,850 for one transport alone, each one probed and keyed. Most store traffic goes to data centres or the internet, so designs use hub-and-spoke, regional meshes or tunnels built on demand between stores that actually talk.
saying these in an interview costs you the question
- SD-WAN removes the need for physical circuits at the branch
- The broadband ISP has to learn the company's private store prefixes
- Running an overlay makes a lossy broadband line as good as MPLS
- SD-WAN is an IETF standard protocol, like OSPF or BGP
- Two circuits in the same duct count as independent paths once tunnelled