A site-to-site IPsec tunnel comes up, traffic flows from site A to site B, but no replies return; what in the SA and SPD state explains it?
answer
- two SAs, two failure points
- compare per-SA counters
- where does the reply leave B
- unknown SPI is audited
- selectors checked after decryption
basics
~20 sEach direction is its own SA, so one can fail alone. Usually B's replies never match its PROTECT entry, or A drops them: their SPI is unknown to A, or their inner headers fall outside the SA's selectors.
solid answer
~50 sA tunnel is two simplex SAs, and the reverse one fails at a distinct point that per-SA counters reveal. If B's outbound SA counter stays flat, the reply never reached IPsec protection: it was routed to a gateway without the SA, its addresses fall outside B's `PROTECT` selectors (or were translated before the lookup), or an earlier `BYPASS` entry sent it in clear, which A then drops. If B's outbound counter climbs but A's inbound counter does not, A is discarding at the SAD lookup because the SPI is unknown, typically stale state after a restart; that is an auditable event and may draw an `INVALID_SPI` notification. If A counts the packets in and still drops them, the decrypted headers fail the selector check, an auditable event that may draw `INVALID_SELECTORS`. Fix the policy so the selectors mirror each other, rather than repeatedly re-establishing SAs.
go deeper
Recall that an IPsec tunnel has a separate SA for each direction, so traffic can work one way and fail the other.
Explain the reply path: B's SPD lookup and outbound SA, then A's SPI lookup in its SAD and its selector check after decryption.
Diagnose with per-SA counters and audit events on both gateways to pin the drop to routing, B's policy, A's SAD or A's selector check.
Set policy standards across peers, such as mirrored selectors, order rules and translation exemptions, so one-way failures are prevented rather than debugged.
## Why one direction can fail alone An IPsec tunnel is not one object. RFC 4301 defines a Security Association as **simplex**, so a two-way tunnel between gateway A (`203.0.113.1`, protecting `10.1.0.0/16`) and gateway B (`198.51.100.1`, protecting `10.2.0.0/16`) consists of an IKE SA plus at least two ESP SAs. An IKE SA that is up proves only that the gateways authenticated each other and negotiated a pair. It does not prove both halves of the pair are being used, or accepted, by both sides. | SA | Outbound at | Inbound at | SPI chosen by | |---|---|---|---| | A to B | A | B | B | | B to A | B | A | A | Traffic from A to B working means the first row is healthy end to end. The fault lives somewhere along the second row. ## Reading the counters Byte-based SA lifetimes mean an implementation already counts traffic per SA, and most expose packet and byte counters; RFC 4301 also makes each IPsec discard below an auditable event. Walk the reply path in order: 1. **Did the reply reach B's IPsec processing at all?** If B's outbound SA counter is flat while users at A are sending, the replies are leaving B some other way. 2. **Did B send ESP but A not accept it?** B's outbound counter rises while A's inbound counter for that SPI stays flat: A is dropping at the SAD lookup. 3. **Did A accept the ESP but discard the contents?** A's inbound counter rises, yet the hosts see nothing: A is dropping after processing, at the selector check. ## Cause 1: the reply never matches B's PROTECT entry B's outbound path consults its SPD before any SA is used, and several things keep a reply off the SA: - **Routing.** RFC 4301 separates forwarding from security decisions; if site B has two exits and the reply is forwarded to a gateway that holds no SA with A, it never meets the tunnel. - **A selector edit on one side.** If B's `PROTECT` entry is changed so its remote side covers only `10.1.1.0/24`, a reply to `10.1.2.5` no longer matches it. RFC 4301 does not require existing SAs to be cleared when the SPD changes, so A's traffic can keep arriving on the old inbound SA at B while B's replies fall to another entry. - **Address translation first.** Where a gateway translates source addresses before its IPsec lookup, the reply's source no longer matches the `PROTECT` selectors. - **Shadowing.** A broad `BYPASS` ordered above the `PROTECT` entry sends the reply out in clear. In each case, a reply that does reach A unprotected is dropped there as plaintext, because A's policy requires that traffic to arrive on an SA. ## Cause 2: A does not know the SPI B stamps each reply with the SPI that A chose when the pair was created. If A has lost that inbound SA, for example after a restart, while B still holds its outbound half, every reply arrives with an SPI absent from A's SAD. RFC 4301 requires A to discard such packets and log an auditable event carrying the SPI. RFC 7296 lets A send an `INVALID_SPI` notification over an existing IKE SA; with none, it may send an unprotected hint that B should treat with suspicion, since it is easy to forge. Recovery means getting a fresh pair established; how that is negotiated is IKE's business. ## Cause 3: A rejects the inner headers After ESP processing, A must match the packet's inner addresses, protocol and ports against the selectors stored with the SA, and discard anything outside them, optionally telling B with `INVALID_SELECTORS`. RFC 7296 describes how this happens with healthy SAs: if A proposes one SA for a whole subnet while its own policy wants one address carried on a different SA with a different algorithm, B may legitimately send traffic from that address on the broad SA, and A drops it. The cure is for A to propose only traffic its own policy accepts on that SA. ## Fixing it - Make the two SPDs mirror images: A's local ranges are B's remote ranges and the reverse, with the same protocols and ports. - Put specific `PROTECT` entries above broad `BYPASS` entries, and exclude tunnel traffic from any translation applied before the IPsec lookup. - Make the reply route lead to the gateway that holds the SA. - Re-establish the SA pair only after the policy is fixed; a new pair negotiated from the same mismatched policy fails the same way.
- How do the audit records tell a lost inbound SA apart from a selector mismatch?An unknown-SPI drop happens before any processing: RFC 4301's audit record carries the SPI, addresses and protocol, and the SPI matches nothing in the SAD. A selector-check drop happens after processing on an SA that exists, and its record adds the selector values from that SAD entry next to the packet's own, which shows exactly which range disagrees.
- Why is an unprotected INVALID_SPI hint treated with suspicion?Without an IKE SA to carry it, the notification travels in clear and anyone can forge it. RFC 7296 says the recipient should treat it only as a hint that something might be wrong and must not respond to it, because responding could cause a message loop.
saying these in an interview costs you the question
- If IKE shows the tunnel up, both directions must be passing traffic.
- One-way traffic always means the keys are wrong.
- An unknown SPI makes the receiver renegotiate automatically.
- Selectors only matter when the SA is first negotiated.
- Re-establishing the tunnel fixes mismatched selectors.