skip to content

AWS creates every Site-to-Site VPN connection with two tunnels. Why two, and what does your side of the connection have to do for that redundancy to be real?

level: seniorimportance: should knowfreq 40%

answer

  1. two endpoints, one of your routers
  2. configure both, not just the one that worked
  3. AWS replaces endpoints on a schedule
  4. BGP withdrawal beats waiting for a timeout
  5. alarm on fewer than two tunnels up

basics

~20 s

AWS terminates each Site-to-Site VPN connection on two independent endpoints so the connection survives the loss or maintenance of one. The redundancy is only real if your customer gateway is configured for both tunnels and your routing fails over — plus a second customer gateway for device redundancy.

solid answer

~50 s

AWS gives each VPN connection two tunnel endpoints, on separate devices with separate outside IP addresses and separate keys, because AWS replaces and patches those endpoints — a single-tunnel deployment takes an outage during routine maintenance. But AWS only supplies half the redundancy. Your customer gateway device must be configured with **both** tunnels up simultaneously; many teams configure one, see traffic flow and stop. Routing decides what failover looks like: with static routing you get one active tunnel and failover only as fast as your device notices, while with BGP both tunnels advertise your prefixes and withdrawal moves traffic in seconds — and on a Transit Gateway you can enable ECMP to use both tunnels at once. Finally, both tunnels of one VPN connection land on the same customer gateway device, so surviving the loss of *your* router needs a second customer gateway and a second VPN connection.

code

bash · 10 lines
bash
# on-prem prefixes learned from the virtual private gateway do not
# reach the VPC route table until propagation is enabled
aws ec2 enable-vgw-route-propagation \
  --route-table-id rtb-0aaa1111 \
  --gateway-id vgw-0bbb2222

# alarm on the silent half-failure: one tunnel down, traffic still flowing
aws cloudwatch describe-alarms-for-metric \
  --namespace AWS/VPN --metric-name TunnelState \
  --dimensions Name=VpnId,Value=vpn-0ccc3333

go deeper

for a junior

Remember that every AWS Site-to-Site VPN connection comes with two tunnels for redundancy, and that your device configuration should bring up both rather than just the first one.

for a middle

Explain what differs between the tunnels — separate AWS endpoints, outside addresses and keys — and how static routing gives one active tunnel while BGP advertises over both and fails over on route withdrawal.

for a senior

Show the operational picture: alarm on the per-tunnel CloudWatch state so a silent single-tunnel failure is caught, verify route propagation into the VPC route tables, and know that AWS replaces endpoints on a maintenance schedule.

for a principal

Own the end-to-end availability target: two customer gateways on independent circuits, four tunnels, ECMP on a Transit Gateway if bandwidth demands it, and a documented failover story between Direct Connect and VPN that has actually been tested.

## Why AWS ships two When you create a Site-to-Site VPN connection, AWS returns two tunnels. Each has its own outside (public) IP address on the AWS side, its own inside addressing, and its own pre-shared key or certificate. They terminate on separate, independently maintained AWS endpoints. The reason is operational: AWS replaces tunnel endpoints for patching, hardware refresh and scaling. Those replacements happen on a schedule and are announced, and if you have only one tunnel configured, each replacement is a hybrid outage. This is the single most common cause of "our VPN dropped and nobody changed anything". ## The half AWS cannot do for you The downloadable device configuration AWS generates contains both tunnels. Bring both up. A surprising number of deployments configure tunnel 1, watch traffic flow, and never touch tunnel 2 — a redundant design that is in fact single-tunnel, and whose second half is discovered only when the first one goes away. Each tunnel carries its own key material and its own liveness detection. Dead peer detection settles what happens when the far side stops answering, and AWS exposes per-tunnel options for the DPD timeout action and for whether AWS or your device initiates the tunnel — the initiation setting matters when your device sits behind NAT and cannot be reached inbound. ## Static routing versus BGP How the two tunnels behave in practice is a routing question. **Static.** You declare your on-premises prefixes on the VPN connection. Traffic uses one tunnel; failover depends entirely on your device detecting the failure and moving traffic, and the AWS side has no way to learn that your network changed. Simple, and slower to converge. **Dynamic (BGP).** Your customer gateway peers over each tunnel and advertises your prefixes; AWS advertises the VPC prefixes back. If a tunnel drops, its session drops and the routes are withdrawn, and traffic moves within seconds. This is also how you add or remove on-premises networks without editing the VPN connection. **ECMP.** Terminating on a Transit Gateway rather than a virtual private gateway lets you enable equal-cost multipath so both tunnels — and tunnels of multiple VPN connections — carry traffic simultaneously. This is the supported way to exceed the per-tunnel throughput ceiling, since individual flows still pin to one tunnel and you are aggregating across flows. ## Getting the routes into the VPC A tunnel that is up but invisible in the VPC route table moves nothing. With a virtual private gateway, you enable route propagation on the VPC route tables so learned on-premises prefixes appear there; otherwise you add static entries pointing at the gateway. With a Transit Gateway, the VPN attachment propagates into the gateway's route tables, and the VPC route tables carry a route to the Transit Gateway. Forgetting propagation produces the exact same symptom as a down tunnel, so check it before you go looking at IPsec. ## The redundancy the two tunnels do not provide Both tunnels of one VPN connection terminate on **one** customer gateway — one device, one internet circuit, quite possibly one building. The two tunnels protect against an AWS-side endpoint failure, not against your router dying. Full redundancy means: - two customer gateway devices, ideally on different internet circuits; - a VPN connection from each, giving four tunnels; - BGP so route withdrawal handles failover between devices. And if a VPN is only the backup for a Direct Connect, remember that both paths should advertise the same prefixes so the failover is automatic rather than a change ticket. ## Watching it AWS publishes a per-tunnel state metric in CloudWatch, so alarming on "fewer than two tunnels up" catches the silent half-failure — the case where one tunnel died, everything still works, and nobody notices until the second one goes. That alarm is the practical payoff of understanding the two-tunnel design, and it is the answer interviewers most want to hear after the mechanics.

  • Both tunnels are up but on-premises traffic still cannot reach instances in the VPC. Where do you look first?
    Routing, not IPsec. With a virtual private gateway, route propagation must be enabled on the VPC route tables — or static routes added — before learned on-premises prefixes appear there. With a Transit Gateway, check that the VPN attachment propagates into the right gateway route table and that the VPC has a route to the gateway.
  • How do you use both tunnels at the same time rather than one active and one standby?
    Terminate the VPN on a Transit Gateway and enable equal-cost multipath. That spreads flows across both tunnels, and across tunnels of multiple VPN connections, which is the supported way to aggregate past the per-tunnel throughput limit. A virtual private gateway does not offer ECMP across tunnels.
  • Do two tunnels make the hybrid link fully redundant?
    No. Both tunnels of one VPN connection terminate on the same customer gateway device, so they protect against an AWS-side endpoint failure only. Surviving the loss of your own router or circuit needs a second customer gateway with its own VPN connection, and BGP so route withdrawal fails traffic over automatically.

saying these in an interview costs you the question

  • One tunnel is enough; the second is just a spare you configure later
  • Both tunnels are always active and load-balance by default
  • Two tunnels mean the on-premises side is fully redundant
  • A tunnel showing UP proves traffic will reach the VPC
  • AWS never takes a tunnel endpoint down

context