How does an IKEv2 gateway decide its peer is dead, and why must it not treat an ICMP unreachable or an unprotected notify as proof?
answer
- only outbound traffic is the warning
- a request with nothing inside
- every IKE request needs a response
- exponential retransmission, then delete
- INITIAL_CONTACT after a reboot
basics
~20 sWhen nothing protected has arrived recently, an IKEv2 gateway sends an empty INFORMATIONAL request and retransmits it with exponential backoff; only repeated silence, or a protected INITIAL_CONTACT on a new IKE SA, proves failure. ICMP and unprotected notifies are forgeable.
solid answer
~50 sRFC 7296 calls it a liveness check (often called dead peer detection). If no cryptographically protected message has arrived on an IKE SA or any of its Child SAs recently — typically when traffic flows only outbound — the gateway sends an `INFORMATIONAL` request whose Encrypted payload holds no payloads, and the peer must answer it. Any fresh protected message proves the IKE SA and all its Child SAs alive. The sender retransmits at exponentially increasing intervals — the RFC suggests at least a dozen times over several minutes, leaving the numbers to implementations — and then deletes the IKE SA and its Child SAs. A peer that rebooted can send `INITIAL_CONTACT` in its first `IKE_AUTH`, letting the other side drop stale SAs at once. An ICMP error or an unprotected notify, such as one about an unknown SPI, may prompt a check but must not end the SA, because anyone can forge one.
go deeper
Recall that IKEv2 checks a quiet peer with an empty INFORMATIONAL request and tears the tunnel down only after repeated silence.
Explain when a check is needed, why any protected message proves liveness, how retransmission backs off, and what INITIAL_CONTACT lets a rebooted peer do.
Separate the RFC's rules from local timer settings, explain why forged ICMP or notifies must not tear tunnels down, and tune detection for failover without inviting false teardowns.
Balance fast failover against false positives on lossy paths, and decide where IKE liveness should be backed by routing-level detection in a redundant tunnel design.
## Why IPsec needs a liveness check An IKE endpoint may lose all its state at any time — a crash, a reboot, a failover to a box that never saw the SAs. Its peer still holds the IKE SA and the Child SAs and keeps encrypting packets towards it. Those packets fall into a **black hole**: they are sent, nothing complains in a trustworthy way, and the tunnel looks up on one side only. RFC 7296 section 2.4 therefore requires that an endpoint confirm the other side is alive when it has reason to doubt it. The RFC notes the mechanism is sometimes called **dead peer detection (DPD)**, although it really detects live peers. ## The mechanism: an empty INFORMATIONAL exchange 1. The gateway notices that **no cryptographically protected message** has arrived on the IKE SA or any of its Child SAs recently. Receiving traffic is itself the proof of life, so the case that matters is a gateway that has only been sending. 2. It sends an `INFORMATIONAL` request containing an IKE header and an Encrypted payload with **no payloads inside**. It is protected with the IKE SA's keys, like every message after setup. 3. The peer MUST respond — every IKE request gets a response, and an empty one is allowed. 4. If no response arrives, the sender **retransmits** the identical request. Retransmission intervals MUST increase exponentially so a congested path is not made worse. 5. Only when repeated attempts have gone unanswered for a timeout period does it conclude the peer has failed, then **discard the IKE SA and every Child SA** negotiated under it. ## What counts as proof | Evidence | May prompt a liveness check | May conclude the peer failed | |---|---|---| | ICMP unreachable or other routing information | yes | no | | Unprotected notify, e.g. about an unknown SPI | yes, rate-limited | no | | Fresh protected message on the IKE SA or a Child SA | — | proves it alive | | Repeated unanswered requests over a timeout period | — | yes | | Protected `INITIAL_CONTACT` on a new IKE SA from the same identity | — | yes, for the old SAs | The rule comes from IKE's design goal of surviving denial-of-service attacks: anything an off-path attacker can forge — ICMP, or a notify sent outside the IKE SA's protection — must never be enough to tear a tunnel down. RFC 7296 also says implementations MUST limit the rate at which they act on unprotected messages, and may ignore them when a protected message arrived recently. ## Timers belong to the implementation RFC 7296 deliberately leaves the number of retries and the timeouts unspecified, because they do not affect interoperability. Its only guidance is a suggestion — retransmit at least a dozen times over at least several minutes — plus the exponential backoff requirement. Any specific "probe every N seconds, give up after M tries" value is a local setting, not the protocol's rule. The same reasoning removes any need to negotiate an SA lifetime: if repeated requests go unacknowledged, the IKE SA and its Child SAs are deleted. The RFC adds a design constraint: because liveness of the IKE SA vouches for all its Child SAs, an implementation must stop sending over any SA if a failure prevents it receiving on all associated SAs, and Child SAs that can fail independently without the IKE SA being able to send a Delete must be negotiated under separate IKE SAs. ## After a reboot: INITIAL_CONTACT A rebooted peer that builds a new IKE SA can include **`INITIAL_CONTACT`** in its first `IKE_AUTH` message. It asserts that this IKE SA is the only one active between the two authenticated identities, so the recipient may delete older IKE SAs to that identity at once instead of waiting for a timeout. It MUST NOT be sent by an identity that may legitimately be connected more than once, such as a roaming user allowed on two devices. ## The IKEv1 equivalent IKEv1 had no built-in liveness exchange. Informational RFC 3706 described a traffic-based method using `R-U-THERE` and `R-U-THERE-ACK` notifies, sent only after an idle period with traffic to send; its idle interval, the "worry metric", was left to each implementation. IKEv1 itself is now deprecated by RFC 9395.
- Why does an IKEv2 gateway that is receiving protected traffic not need to send liveness checks?Because a fresh cryptographically protected message on the IKE SA or any of its Child SAs already proves the peer is alive. RFC 7296 singles out the case where only outgoing traffic has been seen: that is when a gateway could be sending into a black hole without noticing, so that is when a liveness check is needed.
- What must an IKEv2 implementation do if one of its Child SAs can fail while the IKE SA keeps working?RFC 7296 treats liveness of the IKE SA as liveness of all its Child SAs, so the implementation must not break that assumption. It must stop sending on any SA if a failure prevents it receiving on all associated SAs, and Child SAs that can fail independently without the IKE SA being able to send a Delete must be negotiated under separate IKE SAs.
saying these in an interview costs you the question
- An ICMP port unreachable from the peer proves the tunnel is dead.
- RFC 7296 fixes the liveness probe interval and retry count.
- IKEv2 requires liveness probes on a fixed timer even while traffic arrives.
- A liveness check is an unencrypted IKE ping to the peer.
- When the peer stops answering, only its Child SAs are deleted.