skip to content

A caller's SRTP-protected SIP INVITE forks to two phones that both send early media; why can the caller hear silence or clipping, and how do SDES and DTLS-SRTP differ?

level: seniorimportance: should knowfreq 10%

answer

  1. media races the answer
  2. whose key, carried where
  3. one offer, several answers
  4. same port, different keys
  5. a handshake per answerer

basics

~20 s

With SDP security descriptions each callee's sending key arrives only in its own answer, so early media that races ahead cannot be decrypted, and forked streams cannot be matched to keys. DTLS-SRTP keys each callee on its own media-path association.

solid answer

~50 s

Early media flows before the final response, and with forking several early dialogs answer one offer (RFC 3960, Informational). With SDP security descriptions (RFC 4568), each `a=crypto` key is the sender's own key, so the caller learns a callee's key only from that callee's answer; media arriving first has no context and is discarded, so the start is clipped. With forking, every answerer sends to the caller's one transport address with a different key, and RFC 4568 section 7.3 notes the caller cannot associate a packet with an answer; SSRCs may also collide, and every forked phone has seen the caller's key. DTLS-SRTP (RFC 5763) avoids most of this: an answerer wanting early media takes `setup:active` and handshakes at once, each answerer gets its own DTLS association, and the caller matches each certificate to the fingerprint in that answer. Until the answer arrives, though, the fingerprint cannot be checked.

go deeper

for a junior

Recall that early media is ringback or announcements sent before a call is answered, and that forking sends one call to several phones at once.

for a middle

Explain that an a=crypto key is the sender's own key, so a caller cannot decrypt a callee's media until that callee's answer arrives.

for a senior

Diagnose clipped or silent call starts as a key-before-media race, and explain why DTLS-SRTP's per-answerer association and fingerprint check handle forking where SDES cannot.

for a principal

Decide where forking is absorbed, at the border or the endpoint, and whether early media is worth the window in which the peer is not yet authenticated.

## Early media and forking In SIP, a call is not accepted until the callee sends a final `200 (OK)`. **Early media** is media exchanged before that: ringback, announcements, or IVR prompts that a caller answers with DTMF. RFC 3960, an Informational document, describes two ways to run it. In the **gateway model**, early media is negotiated by offer/answer in reliable provisional responses, PRACKs and UPDATEs, inside the early dialog that the INVITE created. In the **application server model**, early media gets its own offer/answer exchange marked with the early-session disposition type, separate from the session that will carry the call. **Forking** is a proxy sending one INVITE to several destinations at once (parallel) or one after another (serial). Each destination that responds creates its own early dialog and its own SDP answer to the caller's single offer. RFC 3960 says the gateway model "is seriously limited in the presence of forking": the caller may get several early media streams on the same address, has to pick one and mute the rest, and may clip speech when a muted dialog is the one that answers. ## Why SDP security descriptions struggle With **SDP security descriptions** (SDES, RFC 4568), the `a=crypto` attribute carries a **master key in the SDP**, and it is the key the *sender of that SDP* will use for its own outgoing media. The caller's offer carries the caller's sending key; each answer carries that callee's sending key. A trace of a forked call: 1. The caller sends an INVITE whose offer contains its key `K_A` and its receive address. 2. The proxy forks to phone B and phone C; both now hold `K_A`. 3. B starts early media to the caller's address, protected with its own key `K_B`, before its answer reaches the caller. 4. The caller has no context for B's SSRC, so the packets are discarded; RFC 4568 notes clipping "may occur until Alice receives Bob's answer". 5. C does the same with `K_C`. When both answers arrive, the caller holds two keys but sees two streams on one port. ## The failure modes RFC 4568 lists - **No association.** Different answerers choose different suites and keys, and "there is no way for the offerer to associate a particular incoming media packet with a particular answer". - **SSRC collision.** Two answerers may pick the same SSRC, which confuses the per-SSRC contexts. - **Key exposure.** Every phone that received the offer knows `K_A`. RFC 4568 says the caller should run a new offer/answer with a new key toward the answerer it keeps. - **Suggested remedy.** For each answer beyond the first, send a new offer with a **new receive address**, turning one-to-many into several one-to-one sessions, and use security preconditions so media does not start before the update is processed. RFC 4568 offers this as one possible approach, not a normative procedure. ## How DTLS-SRTP changes the picture DTLS-SRTP (RFC 5763 and RFC 5764) moves key agreement onto the media path: - **Early media.** An endpoint that wants to send early media MUST take the `setup:active` role and can start the DTLS handshake at once (RFC 5763 section 6.2). - **Forking.** Each answerer forms its own DTLS association with the caller, from its own address, so its packets carry their own keys. The caller "can then securely correlate the SDP answer" by comparing the fingerprint in each answer with the certificate of each association (section 6.3). - **No shared caller key.** No master key travels in the offer, so forked phones learn nothing usable. - **Residual gap.** Until a callee's answer arrives, the caller cannot check that callee's fingerprint, so early media is encrypted but not yet authenticated to a known party. ## What an operator does about it | Symptom | Likely cause | Control | |---|---|---| | First syllables lost on secure calls | SDES key arrives after the media | Security preconditions, or DTLS-SRTP | | Ringback from the wrong phone, or silence | Several early dialogs on one port | Application server model, or one dialog chosen and others muted | | One-way audio after answer | Kept dialog was muted; re-offer not yet processed | Re-offer promptly; expect clipping in the gap | | Forked phones could decrypt the caller | SDES key seen by every fork target | Re-key toward the chosen answerer | A border controller facing WebRTC clients sits in the middle of this. The browser leg expects one answer and DTLS-SRTP, while the carrier leg may fork with SDES, so the border has to collapse the forks into the single dialog the browser sees.

  • Why did RFC 4568 make a=crypto carry the sender's key rather than the receiver's?
    With receive keys, one offer's key would be shared by every forked answerer: SSRC reuse could cause a two-time pad, rollover counters could fall out of step, and no one could track the shared key's lifetime. Sending keys give each answerer its own master key and a unique keystream, at the cost of the caller waiting for each answer before it can decrypt that callee's media.
  • In DTLS-SRTP, what is still unprotected while early media flows before the answer?
    Authentication of the peer. The media is encrypted under keys from a real DTLS handshake, but the caller cannot compare the callee's certificate with a fingerprint it has not yet received, so it does not yet know who it is talking to. RFC 8842 says data received on a DTLS association before the matching fingerprint arrives MUST be treated as coming from an unverified source.

saying these in an interview costs you the question

  • In SDES the caller's a=crypto key protects the media the callee sends.
  • With SDES, forked answers are told apart by their SSRCs, so forking is harmless.
  • SRTP forbids sending any media before the 200 OK arrives.
  • With DTLS-SRTP, every forked answerer shares the caller's master key.
  • DTLS-SRTP early media is authenticated to the callee from its first packet.