skip to content

In WebRTC, how does an SFU use simulcast or SVC to give each receiver a video quality that matches its bandwidth?

level: seniorimportance: should knowfreq 16%

answer

  1. sender offers choices, server picks
  2. independent encodings versus dependent layers
  3. a=simulcast listing rid-ids
  4. a switch waits for a key frame

basics

~20 s

With simulcast the sender encodes the same video at several resolutions as independent streams and the SFU forwards one to each receiver; with SVC it sends one layered stream and the SFU drops upper layers. Neither makes the SFU decode.

solid answer

~40 s

With **simulcast** the sender encodes one camera several times, say 720p, 360p and 180p, configured in the W3C API through `addTransceiver()` with `sendEncodings` entries that each carry a `rid`, and negotiated in SDP with `a=rid` (RFC 8851) and `a=simulcast` (RFC 8853). The SFU estimates each receiver's downlink and display size and forwards one encoding, rewriting RTP headers so the receiver sees a single stream; moving to another encoding waits for a key frame, which the SFU can request. With **SVC** the sender produces one encoding whose temporal or spatial layers build on each other, and the SFU drops upper layers per receiver, which can happen at once. Simulcast costs the sender uplink and encode CPU but works with any codec; SVC needs a codec with scalability support.

go deeper

for a junior

Remember the idea: the sender offers a few qualities and the server hands each viewer the one that fits. Simulcast means several copies; SVC means one stream in layers.

for a middle

Explain where each is configured: sendEncodings with rid in the API, a=rid and a=simulcast in SDP, and that the SFU picks per receiver without decoding.

for a senior

Show the operational details: key-frame waits on simulcast switches, header rewriting for a continuous stream, the uplink cost of extra encodings, and why SVC steps down instantly.

for a principal

Weigh simulcast's universal codec support and sender cost against SVC's single encode and smoother switching, and say how the choice constrains clients and server features.

## One sender, many different receivers In a 12-person meeting the receivers differ: one is on fibre with a large screen, one is on a congested mobile link, one shows the sender as a thumbnail. An SFU cannot serve them all with one encoding without either starving the weak or short-changing the strong. RFC 8853 (§3) names the two ways a middlebox can adapt the view per receiver: **transcode** each stream for each receiver (expensive, adds delay, needs the media content) or **switch** among streams the sender already produced. Simulcast and SVC are the two ways a sender produces streams an SFU can switch among without decoding anything. ## Simulcast: several independent encodings With **simulcast** the sender encodes the same source several times, for example at 720p, 360p and 180p, and sends all of them. Each encoding is an independent RTP stream with its own prediction chain. In SDP, each is described by an `a=rid` line (RFC 8851), which carries restrictions such as `max-width`, `max-height`, `max-fps` or `max-br`, and the `a=simulcast` attribute (RFC 8853) lists the rid-ids sent or received: ``` a=rid:hi send max-width=1280;max-height=720 a=rid:mid send max-width=640;max-height=360 a=rid:lo send max-width=320;max-height=180 a=simulcast:send hi;mid;lo ``` In the W3C WebRTC API, a sender sets this up with `addTransceiver()` and its `sendEncodings` list, each entry carrying a `rid` and, for example, `scaleResolutionDownBy` or `maxBitrate`. The **simulcast envelope** (how many encodings, in what order) is fixed by the first successful negotiation; later renegotiation may narrow it but not re-expand it. `setParameters()` can still pause an encoding by setting `active` to false or change `maxBitrate`. ## SVC: one stream, dependent layers With **SVC** (scalable video coding) the sender produces one encoding made of **layers**: a base layer plus enhancement layers that add frame rate (temporal layers) or resolution (spatial layers). Upper layers depend on lower ones, never the reverse. RFC 8851 (§11.2) notes that scalable layers give an SFU flexibility to forward what best matches each receiver, and that RIDs can express the dependencies with the `depend` restriction. The SFU must be able to tell which layer each packet belongs to, from the codec's payload format or an RTP header extension, but it still never decodes the picture. ## How the SFU chooses per receiver For each receiver and each forwarded source, the SFU repeatedly weighs: - the receiver's estimated downlink bandwidth, from congestion feedback on that leg; - the size the source is displayed at (no 720p for a thumbnail); - which sources matter now, such as the active speaker; - how many sources the receiver wants at all. Moving a receiver to a **different simulcast encoding** is not instant. The new encoding's frames reference earlier frames the receiver never got, so the switch must wait for a **key frame**. RFC 8853 notes it is common to request one with a **Full Intra Request**, and RFC 8834 (§5.1.1) requires WebRTC senders to understand and react to it. The SFU then rewrites RTP header fields so the receiver sees one continuous stream rather than a jump between streams. With **SVC**, stepping down is immediate: drop the upper layers and the lower ones still decode. Stepping up to a higher layer still has to wait for a point where that layer can be decoded. ## Simulcast versus SVC | | Simulcast | SVC | |---|---|---| | What the sender sends | several independent encodings | one layered encoding | | Sender cost | extra encode CPU and uplink for every encoding | one encode; some compression overhead for the layering | | Switching at the SFU | wait for a key frame on the new encoding | drop layers at once; climbing waits for a decodable point | | Codec requirement | any codec | a codec with scalability support | | SDP signalling | `a=rid` plus `a=simulcast` | layer structure in the codec; RIDs can mark dependencies | ## What to remember 1. Neither technique makes the SFU decode; the sender pays to offer choices so the server only has to pick. 2. Simulcast costs the sender's uplink, which is why low layers are kept small. 3. The SFU, not the sender, decides which layer each receiver gets; the sender sends once to the SFU and has no view of each receiver's bandwidth.

  • Why must an SFU wait for a key frame before moving a receiver from the 180p simulcast encoding to the 720p one?
    Each simulcast encoding is independent, with its own prediction chain: its frames reference earlier 720p frames the receiver never received. The receiver can only start decoding at a key frame. The SFU therefore requests one, commonly with a Full Intra Request, which RFC 8834 requires WebRTC senders to understand and react to, and switches at that frame.
  • Can a WebRTC application add a fourth simulcast encoding halfway through a call?
    Not by `setParameters()`. In the W3C WebRTC API the simulcast envelope, the number and order of encodings, is fixed by the first successful negotiation that sends simulcast; later renegotiation may narrow it but not re-expand it. `setParameters()` can still pause an encoding with `active` set to false, resume it, or change `maxBitrate`.

saying these in an interview costs you the question

  • The SFU transcodes the 720p stream down for receivers on slow links.
  • Simulcast saves the sender's uplink compared with sending one stream.
  • With SVC the SFU decodes the base layer in order to strip the others.
  • An SFU can switch a receiver between simulcast encodings on any packet.
  • The sender decides which simulcast encoding each receiver gets.