A site-to-site VPN joins an acquired warehouse whose LAN reuses the data centre's 10.1.0.0/16; why does the tunnel come up while traffic fails, and what fixes it?
answer
- outer addresses versus inner addresses
- the destination looks local
- unambiguous within the VPN
- renumber or present an alias
basics
~20 sThe tunnel forms between the gateways' public addresses, but the inner addresses collide: hosts treat 10.1.0.0/16 as local and never send to the gateway. Renumber one site, or translate so each side sees the other under an unused alias prefix.
solid answer
~50 sA site-to-site tunnel is negotiated between the gateways' outer addresses, which do not collide, so it comes up. The trouble is inside: a warehouse host sending to `10.1.5.10` finds that destination on its own connected subnet and tries to deliver it on the local LAN, so the packet never reaches the gateway; the data centre has the mirror problem, and a gateway cannot route one prefix both onto its LAN and into the tunnel. RFC 4364 states the rule: overlapping address spaces are fine only where those systems never need to communicate. Fixes: **renumber** one site, the durable answer; **translate** at a gateway so each side sees the other under an unused alias prefix, with DNS answering with aliases; or, for a partial overlap, carry only the subnets that do not collide. Separate routing tables keep overlapping ranges apart but do not make them talk.
go deeper
Recall that two sites joined by a VPN need different private ranges, because each host decides whether a destination is local from its own subnet.
Explain outer versus inner addresses: the tunnel uses the gateways' public addresses, so it comes up, while the colliding inner range keeps packets from ever reaching a gateway.
Diagnose the failure quickly, choose between renumbering and two-way translation, and handle the DNS, logging and partial-overlap traps each fix creates.
Make address coordination part of acquisitions and partner onboarding, and set a deadline that turns a translation stopgap into a renumbering plan.
## Why the tunnel is healthy and the traffic is not A site-to-site VPN packet has two IP headers. The **outer** header carries the two gateways' public addresses and is what the internet routes; the **inner** header carries the hosts' own private addresses. The tunnel is set up between the outer addresses, and those are distinct, so authentication succeeds and the tunnel shows as up. The collision is entirely in the inner addresses: the warehouse, bought with its network, numbered its LAN `10.1.0.0/16`, exactly the range the data centre already uses. ## Where a packet goes wrong Follow a warehouse scanner at `10.1.20.7` trying to reach a data-centre server at `10.1.5.10`: 1. The scanner compares the destination with its own subnet, `10.1.0.0/16`, and finds it local. 2. It tries to resolve `10.1.5.10` on the warehouse LAN instead of sending the packet to its default gateway. Either nothing answers, or a different warehouse device that happens to own `10.1.5.10` does. 3. The packet never reaches the warehouse VPN gateway, so it never enters the tunnel. 4. On the data-centre side the mirror image happens for any reply or connection towards `10.1.20.7`. Even configuring the gateways to send `10.1.0.0/16` into the tunnel does not help: each gateway already has that prefix as a connected route on its LAN, and a single routing table cannot hold two meanings for one destination. ## The rule the specifications state - **RFC 1918**, section 3: private addresses are unique only within the enterprise, or within the set of enterprises that choose to cooperate over that space. - **RFC 1918**, section 5: if organisations using private space later interconnect, address uniqueness may be violated; to reduce the risk it strongly recommends choosing private sub-blocks at random. - **RFC 4364**, section 1.3: two VPNs with no sites in common may overlap, which is common with RFC 1918 space, and even VPNs sharing sites may overlap as long as the systems with those addresses never need to communicate; within each VPN every address must be unambiguous. The acquisition has created exactly the case the last rule forbids: two systems with the same address that do need to talk. ## The fixes and what each costs | Fix | How it works | What it costs | |---|---|---| | Renumber one site | move the warehouse to an unused block such as `10.37.0.0/16` | a project: address plans, DHCP scopes, static devices, filters and documentation; afterwards the problem is gone | | Translate at a gateway | the warehouse gateway presents warehouse hosts to the estate as `10.201.0.0/16` and the data centre to warehouse hosts as `10.202.0.0/16`, keeping host bits | DNS must answer with alias addresses on each side; protocols that carry addresses inside their payload can break; logs on the two sides show different addresses for one host | | Carry only non-colliding subnets | if only some subnets overlap, route those that do not | works only if the colliding hosts never need each other | | Separate routing instances | each range lives in its own routing table on a shared device | keeps the ranges apart without letting them communicate; this is how a provider's layer 3 VPN keeps customers apart | Translation has to work in both directions here. If only the warehouse were presented under an alias, warehouse hosts would still find the data centre's real `10.1.x.x` addresses on their own subnet. Each side must see the other under a prefix it does not use itself. The translation mechanics, one-to-one prefix mapping and its state, belong to network address translation; what matters for the VPN design is that each side must reach the other at an unambiguous address. ## Operational traps - **Partial overlaps hide.** A summary such as `10.0.0.0/8` in an address plan can mask a colliding `/24`; compare actual subnets on both sides, including remote-access pools and store ranges. - **Temporary translation becomes permanent.** An alias layer adopted to meet a cut-over date tends to stay, with its DNS views and its confusing logs, so set an end date and renumber behind it. - **Overlap with the tunnel's own pools.** A new site's range can also collide with the remote-access pool or with a store block, failing only for some users. - **Name resolution must agree with routing.** A resolver handing out real addresses across the translation sends hosts straight back into the collision. ## Avoiding the next one - Keep one address registry across every site, pool and partner link. - Choose new private blocks at random, as RFC 1918 recommends, rather than the first range everyone picks. - Make an address-overlap check part of due diligence before connecting an acquired or partner network.
- Why must the translation in this case work in both directions, not just for the warehouse?If only the warehouse is presented to the estate under an alias, data-centre hosts can reach it, but warehouse hosts still see the data centre's real 10.1.x.x addresses as part of their own subnet and deliver locally. Each side must see the other under a prefix it does not use, so the gateway rewrites both source and destination as packets cross.
- Why do separate routing tables on the gateway not solve the overlap?Separate routing instances keep two copies of 10.1.0.0/16 apart, which is how a provider's layer 3 VPN carries customers with the same private ranges. They give each range its own world, but a host in one still cannot address a host in the other, because both are named by the same address. Communication needs distinct addresses, by renumbering or translation.
saying these in an interview costs you the question
- If the tunnel is up, overlapping inner addresses cannot be the problem
- Private addresses cannot be carried inside a site-to-site tunnel
- Translating only the acquired site's addresses lets both sides talk
- Separate routing tables let overlapping sites communicate directly
- A route with a better metric makes the overlapping prefix reachable