skip to content

In a site-to-site VPN, how does store-to-store traffic flow in a hub-and-spoke topology versus a full mesh, and what does each cost?

level: middleimportance: must knowfreq 40%

answer

  1. where the packet detours
  2. decrypt and re-encrypt at the hub
  3. tunnels grow as n(n-1)/2
  4. 302 sites, two hubs

basics

~20 s

In hub-and-spoke, store-to-store traffic detours through a hub, crossing two tunnels and being decrypted and re-encrypted there, but each store keeps only its hub tunnels; a full mesh sends it directly but needs n(n-1)/2 tunnels, 45,451 for 302 sites.

solid answer

~50 s

In **hub-and-spoke**, each store has tunnels only to the hubs, usually the data centres. Traffic from one store to another goes up one tunnel, is decrypted at the hub, routed, re-encrypted and sent down a second tunnel, so it pays a detour in latency, crosses the hub's internet link twice and depends on the hub staying up. In return each store holds two tunnels, adding a store touches only the new store and the hubs, and every inter-store flow passes one inspection point. A **full mesh** gives every pair of sites a direct tunnel: no detour and no hub bottleneck, but `n(n-1)/2` tunnels, so 300 stores and two data centres need 302 × 301 / 2 = 45,451, and every new site means configuration at every other. A **partial mesh** adds direct tunnels only where traffic justifies them.

go deeper

for a junior

Recall the picture: spokes talk to hubs, so store-to-store traffic goes via a hub, while a full mesh gives every pair a direct tunnel.

for a middle

Explain the hub's work per packet, decrypt, route, re-encrypt, and compute tunnel counts: n(n-1)/2 for a full mesh, stores times hubs plus the hub interconnect for hub-and-spoke.

for a senior

Size the hubs for transit traffic counted twice and for one hub carrying everything after a failure, and say where a partial mesh earns its extra tunnels.

for a principal

Weigh path length against central policy and operating cost across the estate, and decide which traffic justifies leaving the hub at all.

## The topology decides the path, not the reachability In a site-to-site VPN each tunnel joins two **VPN edge devices**, the gateways at the sites. RFC 4110, the Informational framework for provider-provisioned layer 3 VPNs, describes the **topology** of a VPN as the set of nodes and the tunnels between them: a **full mesh**, a **hub and spoke** topology, or an arbitrary one. It makes the key point plainly: whatever the topology, all sites remain reachable from each other; the topology only constrains how traffic is routed among them. So the design question is never whether two stores can talk, but which way their packets go and what that costs. ## Hub-and-spoke The stores are **spokes** and the data centres are **hubs**. Every store gateway keeps tunnels to the hubs only. Trace a packet from store 17 to store 42 when both home on data centre A: 1. Store 17's gateway encrypts the packet into its tunnel to data centre A. 2. Data centre A's gateway decrypts it and routes on the inner destination address, finding store 42 behind another tunnel. 3. Data centre A encrypts it again with the keys of the tunnel to store 42. 4. Store 42's gateway decrypts it and delivers it on the store LAN. The hub must decrypt, because each point-to-point tunnel has its own keys, shared only by its two ends. What that detour costs: - **Latency:** the path is store 17 to the hub plus the hub to store 42, not the direct distance. - **Hub bandwidth:** every store-to-store flow enters and leaves the hub's link, so it counts twice against hub capacity. - **Hub processing:** the hub decrypts and re-encrypts every transit packet, so the path costs two encryptions and two decryptions instead of one each. - **Shared fate:** if the only hub fails, store-to-store traffic stops, which is why designs use two hubs. What it buys: - **Few tunnels per store**, one per hub. - **Cheap growth:** a new store needs tunnels to the hubs, and no existing store changes. - **One policy point:** RFC 4110 lists forcing traffic through a firewall, or through a site for monitoring or accounting, as a reason to steer traffic through particular sites. ## Full mesh Every pair of sites has its own tunnel, so a packet between two stores travels directly. RFC 4110 gives the count: with point-to-point tunnels, a full mesh of `N` edge devices needs `N(N-1)/2` duplex tunnels, or `N(N-1)` simplex ones, and each tunnel consumes resources at the edge device. The costs grow with the square of the estate: - every gateway holds `N-1` tunnels and the configuration for each; - a new site needs `N` new tunnels and a change at every existing site; - there is no natural place to inspect traffic between sites. ## The retailer in numbers Take a retailer with 300 stores and two data centres, 302 sites in all. | Design | Tunnels in total | Tunnels per store | Tunnels per data centre | Store-to-store path | |---|---|---|---|---| | Full mesh | 302 × 301 / 2 = 45,451 | 301 | 301 | direct | | One hub | 300 + 1 = 301 | 1 | 301 at the hub, 1 at the other | via the hub | | Two hubs | 300 × 2 + 1 = 601 | 2 | 301 each | via either hub | The `+ 1` is the tunnel between the two data centres. The full mesh is about 75 times the size of the two-hub design, for traffic that in a retail estate mostly flows between stores and data centres anyway. ## Partial mesh: the middle ground RFC 4110 names two reasons for a **partial mesh**: fewer tunnels per device, and a policy need to send certain traffic through a particular site. Its own example maps here directly: a VPN with many telecommuters and a few corporate sites carries mostly telecommuter-to-corporate traffic, so a hub-and-spoke topology fits the traffic, and telecommuter-to-telecommuter traffic still works through a hub. A partial mesh adds direct tunnels where measurement shows heavy, steady traffic, such as between the data centres or among stores sharing a regional warehouse. Some designs go further and build spoke-to-spoke tunnels on demand, using a next-hop resolution step through the hub; the result keeps hub-and-spoke configuration while letting busy pairs take a direct path. ## How to choose - **Most traffic goes to the data centres:** hub-and-spoke, with two hubs for resilience. - **A few pairs talk heavily:** add those pairs as a partial mesh. - **Inter-site traffic is unpredictable and latency-sensitive:** consider on-demand spoke-to-spoke tunnels, and decide where policy is enforced once flows skip the hub. - **A full mesh** suits a handful of sites that all talk to each other, not hundreds.

  • Why must the hub decrypt store-to-store traffic instead of relaying it still encrypted?
    Each point-to-point tunnel has its own keys, shared only by its two ends. A packet arriving from store 17 is protected for the hub, not for store 42, so the hub must decrypt it, route on the inner destination address and encrypt it again for the tunnel to store 42. That puts a decryption and an encryption on the hub for every transit packet, and it is also what lets the hub inspect that traffic.
  • How does the full-mesh count change if tunnels are simplex rather than duplex?
    It doubles. RFC 4110 counts N(N-1)/2 duplex tunnels or N(N-1) simplex tunnels for N edge devices, because a simplex tunnel carries one direction only and each pair needs two. For 302 sites that is 302 × 301 = 90,902 simplex tunnels against 45,451 duplex ones.
  • What happens to store-to-store traffic in a two-hub design when one hub fails?
    Each store still has its tunnel to the surviving hub, so routing moves both store-to-data-centre and store-to-store traffic onto it. Reachability survives, but the surviving hub now carries all transit load, so each hub must be sized to carry the whole estate alone. With a single hub, the same failure stops all store-to-store traffic.

An airline hub: flying between two small cities means two legs through the hub airport. The airline runs far fewer routes and every passenger passes one security checkpoint, but each small-to-small trip takes longer, and the whole network stalls if the hub closes.

saying these in an interview costs you the question

  • The hub relays store-to-store packets without decrypting them
  • In hub-and-spoke, stores cannot reach each other at all
  • A full mesh of n sites needs n squared tunnels
  • A full mesh is always better because every path is direct
  • Adding a site to a full mesh changes only the new site's gateway