In a WebRTC call keyed with DTLS-SRTP, is the audio carried inside DTLS records, and if not, what does the DTLS handshake contribute?
answer
- handshake path versus data path
- use_srtp offers protection profiles
- exporter label EXTRACTOR-dtls_srtp
- client_write and server_write material
basics
~20 sNo. DTLS-SRTP uses DTLS only to authenticate the peers, agree a protection profile through the use_srtp extension and export keying material; RTP and RTCP are then protected by SRTP itself and never sent as DTLS application data.
solid answer
~40 sNo. Once the `use_srtp` extension is negotiated, RFC 5764 says RTP and RTCP are protected solely by SRTP and are never sent in DTLS `application_data` records. The DTLS handshake does three jobs. The client lists the SRTP protection profiles it accepts in `use_srtp` and the server picks exactly one. The certificates exchanged in the handshake authenticate the peers, checked against the SDP `a=fingerprint` lines. After the handshake, a TLS exporter with the label `EXTRACTOR-dtls_srtp` produces a client write master key and salt and a server write master key and salt, and each pair feeds SRTP's own key derivation, so each direction sends under its own keys. Media keeps SRTP's compact per-packet format; DTLS framing on the same port carries only the handshake and other DTLS control messages.
go deeper
Remember the split: the DTLS handshake authenticates the peers and makes keys, and SRTP protects every media packet. Media is not wrapped in DTLS records.
Walk through use_srtp, one profile chosen by the server, then the EXTRACTOR-dtls_srtp exporter split into client and server write keys and salts feeding SRTP's key derivation.
Use the split to diagnose: no use_srtp in the ServerHello means no SRTP keys, and a fingerprint mismatch ends media after ICE succeeded. Explain why a bid-down is caught by Finished.
Explain why the design keeps SRTP's per-packet format and borrows only DTLS key management, and what that buys a media platform: existing SRTP paths, keys never in signalling, no script access.
## Two protocols, two jobs A WebRTC call runs two security protocols on the same UDP port, and they do different work: - **DTLS** (TLS adapted to datagrams) runs a handshake between the two endpoints. It authenticates them with certificates and agrees fresh secret material. - **SRTP** (the Secure Real-time Transport Protocol, RFC 3711) protects each RTP and RTCP packet: it encrypts the payload, appends an authentication tag and lets the receiver reject replays. The combination is **DTLS-SRTP**, defined in RFC 5764, with RFC 5763 describing how it is signalled in SDP. RFC 5764 is explicit: once the `use_srtp` extension is in effect, *application data is never sent in DTLS record-layer `application_data` packets*. Complete RTP or RTCP packets are handed to the SRTP stack, which protects them. Anything that is not RTP or RTCP, such as DTLS handshake messages and alerts, keeps ordinary DTLS framing and travels in separate datagrams. RFC 8827, the WebRTC security architecture, makes DTLS-SRTP mandatory for every media channel: implementations must support SRTP and DTLS-SRTP, must protect all media with SRTP and SRTCP, and must not send plain RTP. ## What the handshake negotiates The client puts a `use_srtp` extension in its ClientHello. Its `UseSRTPData` holds: 1. `SRTPProtectionProfiles`, the profiles the client accepts, in descending order of preference (for example `SRTP_AES128_CM_HMAC_SHA1_80`, which RFC 8827 makes mandatory to support in WebRTC); 2. `srtp_mki`, an optional Master Key Identifier of up to 255 bytes; an empty value means no MKI, and RFC 8827 forbids an MKI in WebRTC. A server that agrees returns its own `use_srtp` in the ServerHello naming **exactly one** profile, and it must not choose one the client did not offer. With no shared profile it should not return the extension, and the connection falls back to an ordinary DTLS cipher suite with no SRTP keys at all. Because the extension is part of the handshake, the Finished messages cover it: a party that strips `use_srtp` to force plain DTLS is detected, which reduces that bid-down attempt to a denial of service. The certificates matter as much as the profile. RFC 5764 notes that the client and server certificates, CertificateRequest and CertificateVerify are all sent in DTLS-SRTP, so both sides are authenticated, and each compares the certificate it received with the `a=fingerprint` its peer put in the SDP. ## From handshake to SRTP master keys After the handshake both sides run a **TLS exporter** (RFC 5705 for DTLS 1.2) with the label `EXTRACTOR-dtls_srtp` and an empty context. It produces `2 * (master_key_len + master_salt_len)` bytes, split in this order: | Slice | Used by | |---|---| | `client_write_SRTP_master_key` | the client to protect what it sends | | `server_write_SRTP_master_key` | the server to protect what it sends | | `client_write_SRTP_master_salt` | paired with the client's key | | `server_write_SRTP_master_salt` | paired with the server's key | With the default profile's 16-octet key and 14-octet salt that is 2 x (16 + 14) = 60 octets. Each key and salt pair goes into one run of SRTP's own key derivation, which turns it into session encryption, authentication and salting keys for SRTP and for SRTCP. The client encrypts with the client write keys, and the server uses those keys only to decrypt and verify what arrives from the client; the mirror holds for the server write keys. Nothing secret crosses the wire: both sides compute the same values from the handshake. ## Why not simply run RTP inside DTLS RFC 5763 gives the reasoning. DTLS's data transfer is generic, while SRTP is tuned for RTP. SRTP leaves the RTP header readable, adds only an authentication tag and an optional MKI to each packet, and reuses SRTP software and hardware that already exists. DTLS-SRTP keeps SRTP's packet format and borrows only DTLS's key management: "the performance benefits of SRTP with the easy key management of DTLS". ## What this means when you troubleshoot - A capture shows DTLS records at the start of the call and SRTP afterwards; that is the design, not a fault. - A handshake that completes **without** `use_srtp` in the ServerHello produced no SRTP keys, so a WebRTC endpoint has nothing to protect media with and the call cannot carry media. - A handshake whose certificate does not match the signalled fingerprint must end the media session, so ICE can succeed and the call can still go silent. - Media keys are never visible in signalling, and RFC 8827 forbids a WebRTC page's script from reading the exported keying material.
- Why does one DTLS-SRTP handshake give the client and the server different SRTP master keys?The exporter output is split into client write and server write master keys and salts. Each side sends under its own write keys and uses the peer's only to decrypt and verify inbound packets. Packets from the two directions therefore never share a key, so an SSRC chosen by one side colliding with the other side's cannot cause keystream reuse across directions.
- What happens if the DTLS server shares none of the SRTP protection profiles the client offered in use_srtp?RFC 5764 says the server should not return `use_srtp`, and the connection falls back to the negotiated DTLS cipher suite without SRTP keys; if that is unacceptable, the server should send a DTLS alert. For a WebRTC call, which may send media only over SRTP, either outcome means no protected media and no audio.
saying these in an interview costs you the question
- WebRTC encrypts the audio packets as DTLS application data records.
- DTLS-SRTP sends the SRTP master key to the peer inside a handshake message.
- Both directions of a DTLS-SRTP call share one SRTP master key.
- The DTLS server may pick an SRTP profile the client never offered.
- Stripping use_srtp from the ClientHello silently downgrades the call to plain DTLS.