Why is a SIP call's media sent as plain RTP over UDP unsafe on a shared office LAN, and what does SRTP add?
answer
- nothing secret, nothing signed
- anyone on the path can listen
- header fields visible, payload type names codec
- RFC 3550 leaves integrity to others
- three services RFC 3711 names as goals
basics
~20 sPlain RTP carries no encryption, integrity check or replay protection: anyone seeing the packets can decode the audio, and anyone reaching the port can inject or replay media. SRTP adds payload encryption, whole-packet authentication and replay protection.
solid answer
~50 sRTP (RFC 3550) is framing, not security: a 12-byte header with `payload type`, `sequence number`, `timestamp` and `SSRC`, then codec bytes. RFC 3550 defines no authentication or integrity service at the RTP level (§9.2), and its optional §9.1 encryption defaults to DES-CBC, which the RFC itself calls too easily broken. So an eavesdropper on the LAN who captures the UDP stream gets audio a decoder plays directly, because the payload type names the codec. A forger needs only the destination port, the SSRC and a sequence number near the current one, all readable in clear, and a recorded packet can simply be sent again. SRTP (RFC 3711) encrypts the payload, authenticates header and payload with a tag, and keeps a replay list so a repeated packet is dropped; with its pre-defined transforms the header itself stays readable.
go deeper
Recall the three gaps of plain RTP — eavesdropping, forgery, replay — and the three services SRTP adds to close them: payload confidentiality, packet integrity, replay protection.
Explain why the gaps exist: RFC 3550 defines no integrity service, the payload type names the codec, and the SSRC and sequence number a forger needs are readable in every packet.
Show you would check both legs: signalling over TLS does not protect media, SRTP without authentication still accepts replays, and voice segregation narrows exposure without authenticating anything.
Frame the trade-off: authenticated SRTP everywhere versus the cost of keying and interop with legacy endpoints, and where a clear header is acceptable metadata exposure for your users.
## What plain RTP is The **Real-time Transport Protocol** (RTP, RFC 3550) carries the audio and video of a call. On a typical office deployment, two SIP phones agree on codecs and ports in the SIP/SDP exchange, then send each other RTP packets over UDP: a phone at `192.0.2.10` sends to `192.0.2.20` on an even port, with **RTCP** (the control companion) on the next odd port. Every RTP packet starts with a **12-byte fixed header**: - `V` (version, always 2), `P` (padding), `X` (extension), `CC` (CSRC count), `M` (marker) - `PT` — the 7-bit **payload type**, naming the codec - a 16-bit **sequence number**, +1 per packet - a 32-bit **timestamp**, the sampling instant of the first octet - a 32-bit **SSRC**, the random identifier of the stream's source After the header come the codec bytes. Nothing in that layout is a secret or a signature. ## What RFC 3550 says about security RFC 3550 is explicit that RTP does not secure itself: 1. **§9.2** — authentication and message integrity "are not defined at the RTP level"; they are expected from lower layers. 2. **§9.1** — an optional confidentiality service encrypts whole packets, by default with DES in CBC mode, which the RFC itself says "has since been found to be too easily broken". 3. **§14** — "an impostor can fake source or destination network addresses, or change the header or payload", and RTCP's CNAME and NAME items can be used to impersonate another participant. So "plain RTP" in practice means media in clear with no check of who sent it. ## The three gaps on a shared LAN | Threat | What the attacker needs | What they gain | Why plain RTP allows it | |---|---|---|---| | **Eavesdropping** | A copy of the packets (on-path or a mirrored port) | Decodable audio | Payload in clear; the payload type names the codec, e.g. PT 0 = PCMU at 8,000 Hz (RFC 3551) | | **Forgery / injection** | The destination port, the SSRC and an in-range sequence number | Speech or noise played into the call | No authentication; the fields it needs are visible in every packet | | **Replay** | A recording of earlier packets | Old audio played again, or a confused receiver | Nothing marks a packet as already seen | RFC 3550 asks for random initial sequence numbers and timestamps, mainly to make known-plaintext attacks on encryption harder, and a random SSRC to avoid collisions. Neither is authentication: an on-path observer reads all three from the first packet. ## What SRTP adds The **Secure Real-time Transport Protocol** (SRTP, RFC 3711) is an RTP profile whose stated goals are: - **confidentiality** of the RTP and RTCP payloads; - **integrity** of the entire RTP and RTCP packets, via an authentication tag over header and payload; - **protection against replayed packets**, via a replay list the receiver keeps over the packet index. These services are optional and independent in RFC 3711, except that integrity is mandatory for SRTCP; RFC 3711 also says SRTP SHOULD NOT be used without message authentication, because encryption alone lets an attacker replay a packet "with certainty that the receiver will accept it". What SRTP does not do with its pre-defined transforms is hide the header: the payload type, sequence number, timestamp and SSRC stay readable, though authenticated. The keys come from a separate exchange — for example SDES in the SDP or a DTLS handshake on the media path — and the ciphers, the replay window and the keying each have their own subject. ## Signalling security is not media security A common gap in deployments: SIP over TLS protects the signalling, hop by hop between SIP elements (RFC 3261), but the RTP flows on its own UDP ports between the media endpoints, or the media relays between them. Unless the session negotiates SRTP (the `RTP/SAVP` profile in SDP), the audio is plain RTP however well the signalling is protected. ## For the operator - Treat a voice segment carrying plain RTP as readable by anyone who can see or mirror its traffic. - Segregating voice traffic narrows who is on the path, but authenticates nothing. - Require SRTP with authentication on every call leg you control, and check that both signalling and media are protected, not just one.
- Don't the random initial sequence number and SSRC already make forging plain RTP hard?Only for a blind sender off the path, and only partly. RFC 3550 asks for random initial sequence numbers and timestamps mainly to make known-plaintext attacks on encryption harder, and a random SSRC to avoid collisions. Anyone on the path reads all three from the first packet, so randomness is no substitute for an authentication tag.
- If the SIP signalling already runs over TLS, is the call's audio protected?No. TLS on SIP protects signalling hop by hop between SIP elements; the RTP media flows on separate UDP ports between the media endpoints and stays plain unless SRTP is negotiated (`RTP/SAVP` in SDP). TLS on signalling matters to SRTP mainly because SDES carries the SRTP keys inside that SDP.
- RFC 3550 does define an encryption service; why not rely on it?Its §9.1 service defaults to DES-CBC, which RFC 3550 itself calls too easily broken, uses the header's sequence number and timestamp as a weak IV, and provides no integrity or replay protection at all. RFC 3550 itself expects SRTP to be the correct choice for many applications.
Plain RTP is a postcard: every handler can read it, and anyone can drop another card in the box signed with your name. SRTP seals the message in a tamper-evident envelope, but the address and postmark on the outside — the RTP header — stay readable so the post can still sort it.
saying these in an interview costs you the question
- Plain RTP is safe on an internal LAN because UDP packets are hard to capture.
- The RTP sequence number and SSRC prove which phone sent the packet.
- If SIP runs over TLS, the call's audio is encrypted as well.
- SRTP encrypts the whole RTP packet, header included.
- Encrypting the payload alone also stops forged and replayed packets.