skip to content

A plain RTP receiver judges packets only by SSRC and sequence number; why does that let an on-path attacker inject or replay audio, and what closes the gap?

level: seniorimportance: should knowfreq 16%

answer

  1. a validity check, not a security check
  2. RFC 3550 appendix A.1
  3. duplicates count as reordered
  4. two in a row means restart
  5. authenticate the index, then check it

basics

~20 s

RTP's sequence check exists to detect loss and source restarts, not attackers: in-range packets, duplicates and resynchronising runs are all accepted, and nothing authenticates the sender. SRTP's authentication tag plus its replay list close the gap.

solid answer

~50 s

RFC 3550 gives a receiver only weak header checks — version 2, a known payload type — and appendix A.1's `update_seq` routine, with typical values `MAX_DROPOUT` 3000, `MAX_MISORDER` 100 and `MIN_SEQUENTIAL` 2. A new SSRC is valid after two sequential packets; a packet slightly ahead is accepted; one fewer than 100 behind is accepted as "duplicate or reordered"; and after a very large jump, a second consecutive packet resynchronises the stream as a restart. An on-path attacker reads the SSRC and current sequence number from any packet, so injected packets just need in-range numbers, and a recorded run of old packets is accepted after one rejection. Payload encryption alone does not help: RFC 3711 says an attacker can replay "with certainty that the receiver will accept it". SRTP with authentication closes it — the tag covers header and payload, and a replay list over the packet index drops repeats.

go deeper

for a junior

Recall that plain RTP checks plausibility, not identity, so forged and replayed packets can be accepted.

for a middle

Walk through RFC 3550's update_seq: the typical constants, why duplicates pass, and how two consecutive out-of-range packets resynchronise the stream.

for a senior

Argue why SRTP needs both the authentication tag and the replay list, why encryption alone fails, and which signs on a plain RTP call suggest injection.

for a principal

Decide where unauthenticated media is still tolerable, such as a closed lab, and what it costs to require authenticated SRTP on every leg.

## What a plain RTP receiver checks RFC 3550 §9.2 defines no authentication or integrity service for RTP, so a receiver has only plausibility checks. Appendix A.1 lists the **header validity checks** — version 2, a known payload type that is not SR or RR, consistent padding, extension and length — and admits that only the first two are always possible, and they "total just a few bits". Beyond that the receiver relies on the **SSRC** and the **sequence number**. Appendix A.1's example routine, `update_seq`, uses three constants it calls "typical values", not mandatory ones: | Constant | Typical value | Meaning at 50 packets per second | |---|---|---| | `MIN_SEQUENTIAL` | 2 | A new SSRC is valid after 2 in-sequence packets | | `MAX_DROPOUT` | 3000 | Accept a forward gap of up to 1 minute | | `MAX_MISORDER` | 100 | Accept packets up to 2 seconds behind | The routine's job is **statistics and restart detection**: counting loss, counting wrap cycles, and recovering when a sender restarts without warning. ## The routine, traced ```pseudocode // RFC 3550 A.1 update_seq, simplified; probation and wrap counting omitted udelta = (seq - max_seq) mod 65536 if udelta < MAX_DROPOUT: // ahead by a permissible gap max_seq = seq return VALID else if udelta <= 65536 - MAX_MISORDER: // a very large jump if seq == bad_seq: // second packet in a row init_seq(seq) // assume the sender restarted return VALID bad_seq = (seq + 1) mod 65536 return INVALID else: // fewer than MAX_MISORDER behind return VALID // duplicate or reordered ``` Three properties matter to a defender: 1. A **duplicate** of a recent packet lands in the last branch and is accepted; RFC 3550 even lets cumulative loss go negative because of duplicates. 2. A **run of old packets** thousands of numbers behind is a "very large jump": the first is rejected, and the second, being consecutive, triggers a resynchronisation. 3. A **brand-new SSRC** becomes valid after `MIN_SEQUENTIAL` in-sequence packets. None of this asks who sent the packet. ## What an attacker needs - **To eavesdrop**: a copy of the traffic. - **To inject**: the destination address and port, which the session's SDP or the traffic shows; the SSRC and a sequence number just ahead of the current one, both readable in every packet; a matching payload type and a plausible timestamp. - **To replay**: a recording of earlier packets from the same session. RFC 3550 §14 adds that an impostor can fake source addresses, so a receiver that checks the source address gains little. Random initial sequence numbers and SSRCs only slow an attacker who cannot see the traffic. Even then the sequence number is a weak barrier: for a known source, `update_seq` accepts about 3,100 of the 65,536 possible values (3,000 ahead, 99 behind), under 5%, and two consecutive packets resynchronise anyway. The 32-bit SSRC is the main obstacle off the path, and no obstacle at all on it. ## What closes the gap RFC 3711 (SRTP) addresses both attacks together: - **Authentication tag.** Computed over the RTP header and the encrypted payload, so a forger without the session key cannot produce an acceptable packet, and a rewritten SSRC or sequence number fails the check. The tag is RECOMMENDED, and RFC 3711 says SRTP SHOULD NOT be used without it. - **Replay list.** The receiver tracks which packet indices it has already accepted and drops a repeat; because the tag authenticates the sequence number, a replayed packet cannot be relabelled with a fresh one. The window size, rollover counter and processing order belong to the replay-protection mechanism itself. Encryption alone is not enough. RFC 3711 §9.5.1 spells out that without message authentication an attacker "can still replay a previous message with certainty that the receiver will accept it", and can flip ciphertext bits when the plaintext is predictable. ## Detecting it on a plain RTP call - Sudden SSRC changes mid-call without a transport change. - Sequence jumps followed by resynchronisation, visible as reset loss statistics. - Negative cumulative loss in RTCP reports, a sign of duplicates. - Audio the speaker did not say. These are clues, not controls; only authentication stops the packets being played.

  • Why can't the receiver simply reject any duplicate sequence number in plain RTP?
    It could reject exact duplicates, but that would not stop forgery: an attacker just uses the next unused number. Without a tag binding the sequence number to the sender, any check on the number is a check on a value the attacker chooses. That is why SRTP's replay list is only meaningful alongside its authentication tag.
  • Would checking that RTP arrives from the address in the SDP stop injection?
    It raises the bar for an off-path sender but does not authenticate anything: RFC 3550 §14 notes that an impostor can fake source addresses, and an on-path attacker sees which address and port to use. RTP itself does not require such a check; any address filter is an implementation choice.

saying these in an interview costs you the question

  • RTP's sequence number check rejects replayed packets.
  • An attacker must guess the 32-bit SSRC before injecting packets.
  • Encrypting the payload alone stops replay of recorded packets.
  • Checking the source IP address authenticates RTP packets.
  • A plain RTP receiver rejects a new SSRC for the rest of the call.