skip to content

Designing media security for a border controller between WebRTC clients and a SIP carrier, when would you terminate and re-originate SRTP rather than relay it end to end?

level: principalimportance: nice to knowfreq 6%

answer

  1. two legs, two rulebooks
  2. who may hold plaintext
  3. keys neither side fully controls
  4. rewriting headers breaks the tag
  5. secure to the border, not beyond

basics

~20 s

Terminate whenever the legs differ: WebRTC must use DTLS-SRTP and may not use SDES, while carriers often accept only plain RTP or SDES-keyed SRTP. Relay end to end only when both ends share keying and the border needs no media access.

solid answer

~40 s

RFC 8827 requires WebRTC endpoints to protect all media with DTLS-SRTP and forbids SDP security descriptions, while a SIP carrier often offers `RTP/AVP` or `RTP/SAVP` with `a=crypto`. When the legs differ, the border must terminate: run its own DTLS handshake toward the browser, decrypt, and re-protect toward the carrier under a separate context, regenerating RTCP per leg. That makes the border a trusted holder of plaintext, costs per-packet crypto, and ends the browser's protection at the border. Relaying end to end works only if both ends speak the same keying and the border leaves RTP headers and payload untouched, since RFC 3711 notes a translator that rewrites them breaks authentication. Choose termination when interworking, transcoding or recording is required; choose relay when confidentiality from the operator matters and both ends can meet.

go deeper

for a junior

Recall that a border controller sits between browsers and a phone carrier, and that it either passes encrypted media through or decrypts and re-encrypts it.

for a middle

Explain why WebRTC's DTLS-SRTP-only rule and a carrier's SDES or plain RTP force the border to terminate media.

for a senior

Show that per-leg contexts, regenerated RTCP and the ban on rewriting authenticated headers follow from the choice, and diagnose silent calls on one leg accordingly.

for a principal

Own the trade-off: interworking, recording and transcoding against holding plaintext for every call, and be explicit about where protection ends.

## The two legs do not speak the same media security A **border controller** here is the element that sits between WebRTC clients and a SIP carrier, carrying both signalling and media. The two sides arrive with different rules: | Aspect | WebRTC leg | SIP carrier leg (common) | |---|---|---| | Media protection | SRTP and SRTCP required, no NULL cipher | plain RTP, or SRTP | | Keying | DTLS-SRTP only; SDP security descriptions forbidden | SDP security descriptions (`a=crypto`) or none | | SDP profile | `UDP/TLS/RTP/SAVPF` | `RTP/AVP` or `RTP/SAVP` | | Ports | usually rtcp-mux and BUNDLE on one transport | often separate RTP and RTCP ports | | Connectivity | ICE with STUN checks | often none | RFC 8827 section 6.5 states the WebRTC side: DTLS-SRTP MUST be offered for every media channel, plain RTP MUST NOT be sent, SDP security descriptions MUST NOT be offered or selected, and an MKI MUST NOT be used. ## Option A: relay end to end The border forwards SRTP without holding keys. For that to work: - Both endpoints must agree on one keying method. A WebRTC client only does DTLS-SRTP, so the far end must too; an SDES-only or plain-RTP carrier rules this out. - The DTLS handshake runs between the two endpoints, and each `a=fingerprint` must reach the other unchanged. - The border must not rewrite anything SRTP authenticates. RFC 3711 section 7.4 notes that mixers and translators that manipulate the RTP header or payload break authentication and end-to-end confidentiality; a trust model may instead let them decrypt and re-encrypt, breaking end-to-end security. - If the relay changes the media endpoint on one side without the other side knowing, RFC 4568 section 7.1.4 says it "cannot operate as a simple packet reflector" and must decrypt and re-encrypt. What the border gives up: no transcoding, no recording, no DTMF interworking, no regenerated RTCP. It can still relay packets and enforce addresses. ## Option B: terminate and re-originate The border is a full endpoint on each leg: 1. Toward the browser it answers with its own fingerprint, completes ICE and DTLS, and derives SRTP keys for that leg. 2. It classifies incoming packets by first byte, decrypts SRTP, and verifies and terminates SRTCP. 3. Toward the carrier it sends a new offer in the carrier's profile, with its own `a=crypto` key over protected signalling, or plain RTP on a private interconnect. 4. It runs separate replay windows, rollover counters and key lifetimes per leg, and produces RTCP reports for each leg from what that leg actually saw. RFC 5763 section 6.9 explains why splicing is impossible: each DTLS handshake "establishes fresh keys that are not completely under the control of either side", so the border cannot hand the browser's key to the carrier. ## What each choice costs | | Relay end to end | Terminate and re-originate | |---|---|---| | Interworking with SDES or plain RTP | impossible | yes | | Plaintext at the border | never | always | | Transcoding, recording, DTMF | no | yes | | Crypto work at the border | none per packet | decrypt and encrypt every packet | | What the browser's "secure" means | secure to the far endpoint | secure to the border only | | Identity assertions | can survive | broken if the SDP is modified (RFC 5763 section 8.6) | ## How to decide - **Interworking need.** A carrier that cannot do DTLS-SRTP forces Option B. - **Trust in the operator.** If users must be protected from the platform itself, Option A, or a double-encryption design such as RFC 8723 for switching conferences, is the only fit. - **Product needs.** Lawful recording, transcoding or announcements need plaintext, so Option B. - **The carrier leg's own risk.** With Option B the carrier leg's protection is whatever it negotiates; plain RTP there means the call is protected only to the border, and the platform should say so. - **Operating cost.** Option B scales with per-call crypto and holds key material for every live call, so the border becomes a high-value target to harden and monitor. Where the carrier cannot do DTLS-SRTP the design is hop by hop by necessity, and the work lies in being explicit about where protection ends, keeping the carrier leg encrypted whenever the carrier supports it, and never presenting a hop-by-hop call as end-to-end secure.

  • Why can't the border just pass the browser's DTLS-SRTP keys to the carrier leg?
    DTLS-SRTP keys come from a handshake whose secrets neither side fully controls (RFC 5763 section 6.9), and RFC 8827 forbids exposing the negotiated keying material to the application. The border can only hold keys for handshakes it takes part in itself, so it must run its own DTLS toward the browser and set up a separate context toward the carrier.
  • If the border terminates SRTP, what should the platform tell users about call security?
    That media is protected between the client and the platform, and on the carrier leg only as far as that leg's own negotiation goes, which may be plain RTP. Calling such a call end-to-end encrypted would be false: the border holds plaintext by design and must be secured and audited as a party to every call.

saying these in an interview costs you the question

  • A border can splice two DTLS-SRTP legs by copying the keys between them.
  • A WebRTC client can fall back to SDES keys if the carrier insists.
  • Rewriting SSRCs and sequence numbers is harmless to opaque SRTP relaying.
  • Terminating SRTP at the border keeps the call end-to-end encrypted.
  • Relaying end to end lets the border still record or transcode the call.