skip to content

After QoS queuing was added after encryption on an IPsec gateway, the peer discards low-priority ESP packets as replays; what in ESP's anti-replay window explains it, and what fixes it?

level: seniorimportance: should knowfreq 22%

answer

  1. numbers assigned before queuing
  2. right edge is the highest validated
  3. left of the window means rejected
  4. one SA per traffic class

basics

~20 s

Sequence numbers are assigned at encryption; priority queuing then lets later numbers overtake. The receiver's window, 64 by default, moves past the delayed packets, which fall left of it and are dropped. Fix: one SA per traffic class, or a larger window.

solid answer

~50 s

An ESP receiver keeps a sliding window whose right edge is the highest sequence number it has validated; anything left of the window, or already marked inside it, is discarded before any cryptography. RFC 4303 requires at least 32 packets and recommends 64 as the default. When the sender numbers packets at encryption and only then queues them by class, a burst of high-priority packets with later numbers can overtake waiting bulk packets; once more than a window's worth have arrived, the bulk packets are too old and are dropped as replays — no attacker involved. RFC 4301 anticipates this: a sender SHOULD put traffic of different DSCP classes on different SAs, each with its own counter and window. The receiver may also enlarge its window locally — it does not tell the sender. Turning anti-replay off removes the protection rather than fixing the ordering.

go deeper

for a junior

Recall that ESP numbers its packets and the receiver keeps a window of recent numbers so it can reject duplicates.

for a middle

Explain the window edges, the 32 minimum and 64 default, and why the window only moves after the ICV passes.

for a senior

Diagnose replay drops caused by post-encryption queuing from counters and captures, then fix them with per-class SAs or a larger window, not by disabling protection.

for a principal

Set a policy for how traffic classes map to SAs across the estate, balancing SA count and rekey load against QoS guarantees.

## The symptom A site-to-site ESP tunnel worked until priority queuing for voice was enabled on the encrypting gateway's egress. Now the far gateway logs anti-replay discards for bulk traffic whenever voice is busy. Integrity failures do not rise, voice is clean, and a capture on the path shows each sequence number once. Nobody is replaying anything; the protocol is reacting to reordering it was designed to treat as suspicious. ## How the ESP anti-replay window works RFC 4303 §3.4.3 describes what the receiver must do, though not how it stores it: - **T** is the highest sequence number that has passed the integrity check — the window's **right edge**. - **W** is the window size. A receiver MUST support at least **32** with 32-bit sequence numbers; **64** is preferred and SHOULD be the default; larger windows MAY be used, and the receiver does **not** tell the sender its size. - The window covers T − W + 1 through T. A number **left** of that is rejected; a number inside it is rejected if already marked as received; a number **right** of T is new. The order matters: 1. A cheap **preliminary check** against the window, before any cryptography, so duplicates are rejected early. 2. **Integrity verification**, or decryption-and-verification under an AEAD cipher. 3. **Window update** — only after the ICV succeeds. 4. Decryption and delivery. ```pseudocode // ESP receiver; W = 64; T = highest validated sequence number on packet(seq): if seq + W <= T: drop("left of window") // too old if seq <= T and seen[seq]: drop("duplicate") if not icv_valid(packet): drop("integrity") // auditable if seq > T: slide window right by (seq - T); T = seq seen[seq] = true decrypt and deliver ``` ## How queuing after encryption causes it The encrypting gateway assigns sequence numbers in the order packets reach ESP processing. Queuing *after* that point changes the transmission order without changing the numbers: 1. Bulk packets 5,001-5,010 are encrypted and wait in the low-priority queue. 2. A burst of voice packets is encrypted as 5,011-5,090 and leaves first from the priority queue. 3. The receiver validates them; T becomes 5,090, so the window now spans 5,027-5,090. 4. The bulk packets arrive. Each satisfies seq + 64 ≤ 5,090, so each is **left of the window** and dropped. The higher the priority load, the further the right edge runs ahead, and the more low-priority packets are lost — exactly the correlation an operator sees. ## Fixes, in order of preference | Fix | Why it works | Cost | |---|---|---| | A separate SA per traffic class | each SA has its own counter and window, so classes cannot overtake each other's numbers; RFC 4301 says a sender SHOULD do this and implementations MUST allow multiple SAs between the same peers | more SAs to negotiate and rekey | | A larger receive window | reordering within the window is tolerated; RFC 4303 says the window should grow for high-speed links | receiver memory; a wider span of old numbers is accepted once | | Classify and queue before ESP processing | numbers are assigned in the order packets will be sent | depends on where the platform allows queuing | | Disable anti-replay on the SA | no drops | replayed packets are accepted: the protection is gone | Disabling anti-replay is the wrong default. RFC 4303 discourages anti-replay for manually keyed SAs and recommends turning it off for multi-sender SAs; it never offers disabling it as a cure for QoS reordering. ## Why the window moves only after the ICV If the receiver advanced T on any packet whose header *claimed* a high number, an attacker could inject one forged packet numbered far ahead and push the window past all genuine traffic, which would then be dropped as old. Updating only after integrity verification means only a packet from the key holder can move the edge, and the preliminary check still rejects obvious replays before spending cycles on cryptography. ## How to confirm the diagnosis - Where the gateway counts discards by reason, they appear as replay or sequence-number failures, not ICV failures. - They hit only the lower-priority classes and track the high-priority load. - A capture shows no duplicated sequence numbers on the wire — just numbers arriving out of order.

  • Why does a larger anti-replay window not require any change at the sending gateway?
    The window is purely receiver state. RFC 4303 lets the receiver pick any size above the minimum and states that it does not notify the sender; the sender's duty is only to number packets in order and never let the counter cycle. Enlarging the window is a local change on the receiving peer, which is why it is often the quickest mitigation while per-class SAs are arranged.
  • Would the same QoS reordering cause drops if the tunnel used AH instead of ESP?
    Yes. AH carries the same 32-bit sequence number and offers the same optional anti-replay window at the receiver's discretion, and RFC 4301's advice about putting different traffic classes on different SAs names both protocols. The cause is reordering after numbering, not anything specific to ESP's encryption.

saying these in an interview costs you the question

  • Replay discards after enabling QoS mean someone is replaying our traffic.
  • The right fix is to turn anti-replay off on the tunnel.
  • The peers must agree the window size during SA negotiation.
  • The receiver advances its window before checking the ICV.
  • ESP sequence numbers are encrypted, so reordering cannot be seen on the path.