In a WebRTC call through an SFU, why is DTLS-SRTP not end-to-end encryption, and how does SFrame (RFC 9605) add it?
answer
- who holds the SRTP keys
- hop by hop, not end to end
- encrypt the frame, not the packet
- keys the server never sees
basics
~20 sDTLS-SRTP keys each hop separately, so an SFU decrypts every packet and can read the media. SFrame encrypts each encoded frame with keys the SFU never holds, inside the hop-by-hop SRTP, leaving the SFU only the metadata it needs to forward.
solid answer
~40 sEvery WebRTC peer connection runs its own DTLS handshake, so with an SFU there are separate DTLS-SRTP associations from sender to SFU and from SFU to each receiver. The SFU decrypts each packet and re-encrypts it for each receiver, so it holds the plaintext encoded media: encrypted on every wire, but not end to end. **SFrame** (RFC 9605) adds a second layer: the sender encrypts each encoded frame, or each packet payload, with an AEAD cipher suite under keys from an end-to-end key-management system, such as sender keys over a secure channel or MLS, and prefixes an SFrame header with a key ID and counter. The SFU still reads RTP headers and the SFrame header and keeps forwarding, but cannot decrypt. SFrame leaves key distribution to the application and gives no per-sender authentication.
go deeper
Know that WebRTC media is always encrypted, but that a server in the middle decrypts each hop. End-to-end means only the participants hold the keys.
Explain why separate peer connections mean separate DTLS-SRTP keys, and that SFrame encrypts frames inside SRTP while leaving headers readable.
Discuss what SFrame leaves to you: key distribution via sender keys or MLS, rotation on join or leave and its key-frame effect, and the lack of per-sender authentication.
Weigh E2EE against server-side features it removes, recording, transcoding, mixing, and decide who may hold keys and how compliance needs are met.
## Hop-by-hop and end-to-end are different promises **Hop-by-hop (HBH) encryption** protects media on each network leg; whoever terminates a leg can read what it carries. **End-to-end encryption (E2EE)** protects media from the sender's encoder to the receivers' decoders, so no server in between can read it. RFC 9605, which defines SFrame, opens with exactly this split: an SFU usually needs RTP metadata and RTCP feedback to do its job, so a conference needs **two layers** — HBH protection of media, metadata and feedback between endpoints and the SFU, and E2E protection of the media itself. ## Why DTLS-SRTP stops at the SFU WebRTC media is always encrypted: RFC 8827 requires SRTP keyed by DTLS-SRTP on every media channel and forbids unencrypted RTP. But the DTLS handshake runs between the two ends of **one** peer connection. With an SFU in the middle there are separate connections — sender to SFU, SFU to each receiver — each with its own keys (how DTLS produces those keys belongs to the SRTP keying material). The SFU decrypts every packet arriving from the sender and re-encrypts it for each receiver, so it holds **plaintext encoded media**. The media is encrypted on every wire, yet any SFU operator, or attacker on the SFU, can read it. In a mesh this problem does not arise, because each DTLS-SRTP association runs directly between two participants. ## What SFrame adds SFrame (Secure Frame, RFC 9605, Standards Track, August 2024) is an encryption framing designed to sit **inside** an HBH-protected transport such as SRTP. The sender protects media before it reaches the transport: 1. **Per frame**: encrypt each whole encoded frame, then packetize the ciphertext — the lowest overhead, but packetization must not depend on the frame's content. 2. **Per packet**: encrypt each media payload after packetization — simpler integration, more overhead. An SFrame ciphertext is an **SFrame header** followed by AEAD-encrypted data, with the header used as additional authenticated data. The header carries a **key ID (KID)** and a **counter (CTR)**; the nonce is the per-key salt XORed with the counter. RFC 9605 defines AES-CTR-with-HMAC suites (such as `AES_128_CTR_HMAC_SHA256_80`) and `AES_128_GCM_SHA256_128`. The SFU still reads the RTP header, any header extensions the application exposes, and the SFrame header, so it can keep forwarding, switching streams and dropping layers. With SVC, the sender **MUST** put each layer in its own SFrame ciphertext so the SFU can drop layers it cannot read. ## The keys: SFrame's deliberate gap SFrame does not say how keys reach participants. RFC 9605 (§5) leaves key management to the application and describes two patterns: - **Sender keys**: each participant generates a base key and sends it to the others over an existing end-to-end-secure channel; the KID encodes a key generation and a ratchet step. - **MLS**: the Messaging Layer Security group protocol gives every member a new shared secret per epoch, from which SFrame keys are exported. The key-management system **MUST** ensure each key encrypts media for exactly one sender, so nonces never repeat, and every client **SHOULD** change keys when someone joins or leaves, for forward secrecy and post-compromise security. The point for a design: the SFU and the signalling server must never be able to obtain these keys, or the E2EE is only nominal. ## What SFrame does not give you | Property | SFrame's position | |---|---| | Media confidentiality from the SFU | yes | | Confidentiality of the SFrame header | no; KID and CTR are visible, only integrity-protected | | Proof of which participant sent a frame | no; keys are symmetric, so any receiver could forge (needs digital signatures) | | Key distribution and rotation | left to the application | | Replay protection | out of scope; the application decides | | Transcoding or mixing at a server | impossible by design; the server cannot decode | Two operational consequences are worth stating in an interview: - **Key rotation versus key frames.** When a participant joins, keys rotate, and the SFU asks senders for a key frame. If that key frame is encrypted under a key the joiner does not hold yet, it is discarded, and video waits for the next key frame. RFC 9605 (§6.2) advises sending a key frame after the new key is in use. - **No server-side processing of content.** Recording, server-side mixing and transcoding stop working unless a component that holds keys, such as a recording client, joins as a participant. ## Why not encrypt SRTP twice? An earlier IETF scheme, SRTP "double encryption" (RFC 8723), adds an inner E2E layer inside SRTP. RFC 9605 judges it to have poor efficiency and high complexity, and its entanglement with RTP makes it unworkable in several realistic SFU scenarios. SFrame is transport-agnostic and aims for minimal packet expansion and minimal SFU changes.
- Why can SFrame not stop one meeting participant from impersonating another?SFrame uses symmetric keys: every receiver of a sender's media holds the key that encrypts it, so any receiver could produce frames that appear to come from that sender. RFC 9605 states it gives no per-sender authentication; preventing impersonation by a malicious participant needs an additional mechanism based on digital signatures.
- Why might a participant who joins an SFrame-protected call see frozen video for a moment?Keys SHOULD rotate when someone joins. The SFU asks senders for a key frame for the joiner, but if that key frame is encrypted under a key the joiner does not hold yet, it is discarded; later frames decrypt but cannot be decoded until the next key frame. RFC 9605 advises sending a key frame once the new key is in use.
- Does SFrame stop an SFU from using simulcast or SVC?No. Each simulcast encoding produces its own frames, each encrypted with a unique counter, and the SFU still selects among them by RTP metadata. For SVC, RFC 9605 requires the sender to put each layer in a separate SFrame ciphertext, so the SFU can drop layers without reading them.
saying these in an interview costs you the question
- DTLS-SRTP is end to end, so an SFU can never see the media.
- SFrame replaces SRTP, so the hop-by-hop encryption can be dropped.
- SFrame defines how participants exchange and rotate their keys.
- With SFrame, every receiver can prove which participant sent a frame.
- An SFU can still transcode SFrame-protected video for slow receivers.