A retailer's dual-hub site-to-site VPN now carries heavy store-to-store video traffic; how do you choose between bigger hubs, regional hubs, a partial mesh and on-demand spoke-to-spoke tunnels?
answer
- measure before redesigning
- transit counts twice at the hub
- configuration debt per static pair
- inspection moves with the path
basics
~20 sStart from the traffic matrix: modest store-to-store volume means bigger hubs; regional clusters mean regional hubs or a partial mesh; unpredictable pairs mean on-demand spoke-to-spoke tunnels, and any flow that skips a hub needs its inspection moved to the stores.
solid answer
~40 sMeasure first: which store pairs talk, how much, at what hours, and whether the traffic is latency-sensitive. Hub transit costs twice the hub's link bandwidth per flow, an extra decrypt and re-encrypt, and a detour; if that is affordable, **bigger hubs** keep the simplest design and one inspection point. If traffic clusters by region, **regional hubs** shorten the detour and spread load, at the price of more hubs to run, secure and keep redundant. A **partial mesh** of static tunnels suits a few heavy, stable pairs, but every pair becomes configuration to maintain. **On-demand spoke-to-spoke** tunnels follow the traffic while keeping hub-and-spoke configuration, but direct flows bypass the hub's inspection and stores must be able to reach each other. A full mesh, 45,451 tunnels for 302 sites, is not a candidate.
go deeper
Recall that store-to-store traffic in hub-and-spoke goes through a hub, and that a direct path is the alternative.
Explain what a hub detour costs per flow, link capacity twice, an extra decrypt and re-encrypt, and added delay, and what a static partial mesh adds.
Size hubs for transit and failover, and describe how on-demand tunnels keep the hub in the control path while moving data off it.
Own the trade-off: measured traffic against inspection, operating cost and blast radius, a recommendation with its trigger for revisiting, and who signs for the policy that moves to the stores.
## Start from the traffic matrix A topology change is a bet on where traffic goes, so the first deliverable is a **traffic matrix**: for each pair of sites, how much traffic, at what times, and of what kind. For a retailer with 300 stores and two data centres the usual finding is that most volume runs between stores and data centres, and store-to-store traffic is a minority concentrated in a few patterns: stores sharing a regional warehouse, video calls between neighbouring stores, a training stream. The decision turns on three numbers: - **volume:** what share of hub load is transit between stores; - **shape:** whether the talking pairs are few and stable, regional, or scattered; - **sensitivity:** whether the traffic is interactive, so the detour's added delay matters, or bulk, so only throughput matters. ## What hub transit actually costs In hub-and-spoke a store-to-store packet goes up one tunnel, is decrypted, routed and re-encrypted at the hub, and goes down another. Each flow therefore costs: - **twice the hub link:** 200 Mb/s of store-to-store traffic adds 400 Mb/s of load on the hubs' links, 200 in and 200 out, on top of store-to-data-centre traffic; - **an extra decryption and encryption** at the hub for every packet; - **added delay** of the path via the hub instead of the direct path; - **shared fate** with the hub, which the second hub only partly offsets, since each hub must be able to carry everything alone. Against that, the hub is the one place every inter-store flow can be inspected, logged and accounted for, and RFC 4110 names that policy need as a reason to steer traffic through particular sites. ## The four options | Option | What changes | Gains | Costs | Fits when | |---|---|---|---|---| | Bigger hubs | capacity only | simplest, one policy point | the detour remains | transit is modest or bulk | | Regional hubs | stores home on a nearby hub, regional hubs link to the data centres | shorter detour, spread load | more hubs to run, secure and keep redundant | traffic clusters by region | | Partial mesh | static tunnels for chosen pairs | direct path for heavy, stable pairs | each pair is configuration to maintain and inspect at both ends | few pairs, predictable volume | | On-demand spoke-to-spoke | tunnels built when traffic appears, torn down when idle | direct paths without per-pair configuration | direct flows skip the hub, stores must reach each other, more moving parts | pairs are many and unpredictable | Two details of the on-demand option matter for the decision. Stores learn where to build a direct tunnel through a resolution step served by the hub, so the hubs stay in the control path. And while that resolution is pending, the next-hop resolution protocol (RFC 2332) recommends forwarding packets along the routed path, here through the hub, so the first packets towards a store with no direct tunnel yet still take the detour. The mechanism belongs to multipoint tunnel design; the design outcome is a hub-and-spoke configuration with mesh-like paths for busy pairs. A full mesh is not on the list: RFC 4110's count of `N(N-1)/2` duplex tunnels gives 302 × 301 / 2 = 45,451, every new store means a change at every site, and almost none of those tunnels would carry steady traffic. ## Policy moves with the path The option that removes the detour also removes the hub from the flow. Before any direct store-to-store path goes live, decide: 1. **Where inter-store policy is enforced.** If not at the hub, then at each store gateway, which means 300 policy points to keep consistent instead of two. 2. **Where the flows are logged.** Evidence that used to come from two hubs now comes from every store. 3. **Whether stores may talk directly at all.** A store that is compromised can now reach other stores without crossing an inspected point; for some estates that ends the discussion and the answer is bigger or regional hubs. ## A defensible recommendation for the retailer 1. Build the traffic matrix over a representative month, including peak trading days. 2. If store-to-store transit is a small share of hub capacity, scale the hubs, sizing each to carry the whole estate after a failure, and stop there. 3. If traffic clusters around regional warehouses, add regional hubs or a small partial mesh for those clusters, keeping inspection at the regional point. 4. Only if pairs are many, scattered and latency-sensitive, adopt on-demand spoke-to-spoke tunnels, together with the store-level policy and logging that direct paths require. 5. Revisit when the matrix changes; a topology chosen for this year's traffic is a hypothesis, not a fixed asset. ## Traps - **Redesigning without data:** the loudest complaint is rarely the largest flow. - **Mesh by accident:** static pairs added one ticket at a time become an unmanaged partial mesh nobody can draw. - **Treating the hub as only a cost:** removing it removes the inspection point too.
- Why do on-demand spoke-to-spoke tunnels still depend on the hubs?Stores learn where to build a direct tunnel through a resolution step the hub serves, so the hubs remain in the control path, and the next-hop resolution protocol's recommended default forwards packets along the routed path, via the hub, until the direct path is resolved. Lose the hubs and stores cannot set up new direct tunnels, so hub redundancy still matters.
- When is a static partial mesh better than on-demand tunnels?When a few pairs carry heavy, steady traffic, such as stores sharing a regional warehouse or the two data centres. Static tunnels can be capacity-planned, inspected at both ends and monitored like any other link, and they avoid the extra moving parts of on-demand setup. On-demand tunnels earn their complexity only when the talking pairs are many and change often.
- What must change in security operations when stores start talking directly?Inter-store policy moves from two hubs to every store gateway, so rule sets must be distributed and kept consistent across 300 sites, and flow logs must be collected from all of them. The team also has to accept that a compromised store can reach peers without crossing an inspected point, or keep those flows on the hubs.
saying these in an interview costs you the question
- A full mesh is the natural fix once stores talk to each other
- Store-to-store traffic costs the hub its volume once, not twice
- On-demand spoke-to-spoke tunnels make the hubs unnecessary
- Direct store-to-store tunnels keep all flows inside the hub's inspection
- The topology should be chosen before measuring traffic between sites