A plain RTP receiver judges packets only by SSRC and sequence number; why does that let an on-path attacker inject or replay audio, and what closes the gap?
answer
- a validity check, not a security check
- RFC 3550 appendix A.1
- duplicates count as reordered
- two in a row means restart
- authenticate the index, then check it
basics
~20 sRTP's sequence check exists to detect loss and source restarts, not attackers: in-range packets, duplicates and resynchronising runs are all accepted, and nothing authenticates the sender. SRTP's authentication tag plus its replay list close the gap.
solid answer
~50 sRFC 3550 gives a receiver only weak header checks — version 2, a known payload type — and appendix A.1's `update_seq` routine, with typical values `MAX_DROPOUT` 3000, `MAX_MISORDER` 100 and `MIN_SEQUENTIAL` 2. A new SSRC is valid after two sequential packets; a packet slightly ahead is accepted; one fewer than 100 behind is accepted as "duplicate or reordered"; and after a very large jump, a second consecutive packet resynchronises the stream as a restart. An on-path attacker reads the SSRC and current sequence number from any packet, so injected packets just need in-range numbers, and a recorded run of old packets is accepted after one rejection. Payload encryption alone does not help: RFC 3711 says an attacker can replay "with certainty that the receiver will accept it". SRTP with authentication closes it — the tag covers header and payload, and a replay list over the packet index drops repeats.
go deeper
Recall that plain RTP checks plausibility, not identity, so forged and replayed packets can be accepted.
Walk through RFC 3550's update_seq: the typical constants, why duplicates pass, and how two consecutive out-of-range packets resynchronise the stream.
Argue why SRTP needs both the authentication tag and the replay list, why encryption alone fails, and which signs on a plain RTP call suggest injection.
Decide where unauthenticated media is still tolerable, such as a closed lab, and what it costs to require authenticated SRTP on every leg.
## What a plain RTP receiver checks RFC 3550 §9.2 defines no authentication or integrity service for RTP, so a receiver has only plausibility checks. Appendix A.1 lists the **header validity checks** — version 2, a known payload type that is not SR or RR, consistent padding, extension and length — and admits that only the first two are always possible, and they "total just a few bits". Beyond that the receiver relies on the **SSRC** and the **sequence number**. Appendix A.1's example routine, `update_seq`, uses three constants it calls "typical values", not mandatory ones: | Constant | Typical value | Meaning at 50 packets per second | |---|---|---| | `MIN_SEQUENTIAL` | 2 | A new SSRC is valid after 2 in-sequence packets | | `MAX_DROPOUT` | 3000 | Accept a forward gap of up to 1 minute | | `MAX_MISORDER` | 100 | Accept packets up to 2 seconds behind | The routine's job is **statistics and restart detection**: counting loss, counting wrap cycles, and recovering when a sender restarts without warning. ## The routine, traced ```pseudocode // RFC 3550 A.1 update_seq, simplified; probation and wrap counting omitted udelta = (seq - max_seq) mod 65536 if udelta < MAX_DROPOUT: // ahead by a permissible gap max_seq = seq return VALID else if udelta <= 65536 - MAX_MISORDER: // a very large jump if seq == bad_seq: // second packet in a row init_seq(seq) // assume the sender restarted return VALID bad_seq = (seq + 1) mod 65536 return INVALID else: // fewer than MAX_MISORDER behind return VALID // duplicate or reordered ``` Three properties matter to a defender: 1. A **duplicate** of a recent packet lands in the last branch and is accepted; RFC 3550 even lets cumulative loss go negative because of duplicates. 2. A **run of old packets** thousands of numbers behind is a "very large jump": the first is rejected, and the second, being consecutive, triggers a resynchronisation. 3. A **brand-new SSRC** becomes valid after `MIN_SEQUENTIAL` in-sequence packets. None of this asks who sent the packet. ## What an attacker needs - **To eavesdrop**: a copy of the traffic. - **To inject**: the destination address and port, which the session's SDP or the traffic shows; the SSRC and a sequence number just ahead of the current one, both readable in every packet; a matching payload type and a plausible timestamp. - **To replay**: a recording of earlier packets from the same session. RFC 3550 §14 adds that an impostor can fake source addresses, so a receiver that checks the source address gains little. Random initial sequence numbers and SSRCs only slow an attacker who cannot see the traffic. Even then the sequence number is a weak barrier: for a known source, `update_seq` accepts about 3,100 of the 65,536 possible values (3,000 ahead, 99 behind), under 5%, and two consecutive packets resynchronise anyway. The 32-bit SSRC is the main obstacle off the path, and no obstacle at all on it. ## What closes the gap RFC 3711 (SRTP) addresses both attacks together: - **Authentication tag.** Computed over the RTP header and the encrypted payload, so a forger without the session key cannot produce an acceptable packet, and a rewritten SSRC or sequence number fails the check. The tag is RECOMMENDED, and RFC 3711 says SRTP SHOULD NOT be used without it. - **Replay list.** The receiver tracks which packet indices it has already accepted and drops a repeat; because the tag authenticates the sequence number, a replayed packet cannot be relabelled with a fresh one. The window size, rollover counter and processing order belong to the replay-protection mechanism itself. Encryption alone is not enough. RFC 3711 §9.5.1 spells out that without message authentication an attacker "can still replay a previous message with certainty that the receiver will accept it", and can flip ciphertext bits when the plaintext is predictable. ## Detecting it on a plain RTP call - Sudden SSRC changes mid-call without a transport change. - Sequence jumps followed by resynchronisation, visible as reset loss statistics. - Negative cumulative loss in RTCP reports, a sign of duplicates. - Audio the speaker did not say. These are clues, not controls; only authentication stops the packets being played.
- Why can't the receiver simply reject any duplicate sequence number in plain RTP?It could reject exact duplicates, but that would not stop forgery: an attacker just uses the next unused number. Without a tag binding the sequence number to the sender, any check on the number is a check on a value the attacker chooses. That is why SRTP's replay list is only meaningful alongside its authentication tag.
- Would checking that RTP arrives from the address in the SDP stop injection?It raises the bar for an off-path sender but does not authenticate anything: RFC 3550 §14 notes that an impostor can fake source addresses, and an on-path attacker sees which address and port to use. RTP itself does not require such a check; any address filter is an implementation choice.
saying these in an interview costs you the question
- RTP's sequence number check rejects replayed packets.
- An attacker must guess the 32-bit SSRC before injecting packets.
- Encrypting the payload alone stops replay of recorded packets.
- Checking the source IP address authenticates RTP packets.
- A plain RTP receiver rejects a new SSRC for the rest of the call.