When would you deliver a 10,000-viewer live session over WebRTC through SFUs rather than HLS or DASH, and at what cost?
answer
- seconds versus under a second
- stateless GET versus stateful session
- who can cache the bytes
- per-viewer encryption and feedback
- what the viewers must do
basics
~20 sChoose WebRTC when viewers must react within a second, as in live bidding or on-camera questions. HLS or DASH take seconds but scale on cacheable HTTP; WebRTC needs a stateful, separately encrypted session per viewer on SFU capacity.
solid answer
~50 sHLS and DASH deliver segments over ordinary HTTP GET, every viewer gets the same bytes, so caches and CDNs serve 10,000 viewers cheaply, but the player buffers whole segments and delay is measured in seconds. WebRTC pushes RTP over UDP with a small jitter buffer, so delay typically stays under a second, but every viewer is a separate peer connection to an SFU: ICE checks and a DTLS handshake to join, SRTP encryption per viewer because RFC 8827 forbids unencrypted media and keys are per connection, plus feedback and congestion control per viewer. Nothing can be cached, so 10,000 viewers means a tree of SFUs and egress that grows with the audience. I pick WebRTC when viewers interact, such as auctions or live Q&A, HLS for passive watching, and often a hybrid: the stage on WebRTC, the audience on HLS.
go deeper
Remember the headline: WebRTC is under a second, HLS and DASH take seconds, and plain HTTP delivery is much easier to scale to a big audience.
Explain where the delay comes from in segmented delivery and what each WebRTC viewer costs: a handshake, its own encryption and its own feedback.
Lay out the fan-out tree of SFUs, the per-viewer egress and encryption bill, join surges, and the hybrid in which the stage is WebRTC and the audience HLS.
Decide by interaction need rather than audience size, and price both paths so the product team sees what sub-second delay costs per viewer.
## Two delivery models for the same live video A live session with 10,000 viewers can reach them in two quite different ways. | | Segmented HTTP (HLS, MPEG-DASH) | WebRTC through SFUs | |---|---|---| | Transport | HTTP GET of segment files | RTP pushed over UDP inside SRTP | | Typical delay | seconds | under a second | | Bytes per viewer | identical, so cacheable | encrypted per viewer | | Server state per viewer | none beyond a request | ICE, DTLS, SRTP and congestion state | | Scales with | ordinary HTTP caches and CDNs | SFU capacity, often in a tree | | Viewer can talk back | no, needs a separate channel | yes, the same session can send | ## Why segmented delivery takes seconds and scales so easily HLS and DASH cut the stream into short segments listed in a manifest. A player fetches segments with ordinary HTTP requests and keeps a few in its buffer before playing. That buffer is where the seconds of delay come from; the low-latency chunked variants shrink it, but they still sit in the range of seconds rather than below one. In return, every viewer receives **the same bytes**, so any HTTP cache can serve thousands of viewers from one copy. The origin barely notices viewer 9,999. How segments, bitrate ladders and caches are tuned is its own subject; for this choice only the contrast matters. ## What WebRTC costs per viewer WebRTC sends media as RTP packets the moment they are encoded, and the receiver keeps only a small jitter buffer, so glass-to-glass delay typically stays below a second. The price is that each viewer is a **separate peer connection** to an SFU: 1. **Connection setup.** Signalling, ICE connectivity checks and a DTLS handshake for each viewer, so joining takes noticeable time and a surge of joins is a surge of handshakes. 2. **Per-viewer encryption.** RFC 8827 makes SRTP mandatory and forbids unencrypted media, and keys are per connection. The SFU encrypts every packet separately for every viewer, so no cache can share the bytes. 3. **Per-viewer feedback.** Congestion control, retransmission requests and key-frame requests arrive from every viewer; a key-frame request from one viewer can ask the sender for a new key frame that every viewer then receives. 4. **Per-viewer adaptation.** Viewers on poor links need lower simulcast encodings or SVC layers, chosen by the SFU per viewer. ## Fanning out to 10,000 viewers One SFU cannot hold every viewer, so large WebRTC audiences are served by a **tree**: the broadcaster sends once to an origin SFU, which forwards to edge SFUs, each of which serves a share of the viewers, often in the viewers' regions. Ingest into such a system can itself be WebRTC; WHIP (RFC 9725) defines an HTTP POST that delivers a broadcaster's SDP offer to a media server and returns its answer. The WebRTC specifications define each endpoint-to-server leg; how SFUs talk to each other in the tree is left to each implementation. The bill grows with viewers in a way it does not for HLS: SFU egress, encryption CPU and connection state for every single viewer, with no shared cache to absorb repeats. ## Deciding Choose by what the viewers must *do*, not by how many there are: - **WebRTC** when viewers act on what they see within a second: live auctions and bidding, betting, a host taking questions from viewers on camera, interactive classes, watch-together sessions. - **HLS or DASH** when viewers mainly watch: a conference keynote, a concert, a sports stream where a few seconds of delay changes nothing. - **Hybrid** for many real products: the people on stage talk over a WebRTC SFU at sub-second delay; the server composes their output and publishes it as HLS or DASH for the large passive audience. A viewer who is invited on stage leaves the HLS player and joins the WebRTC session. ## A common mistake "WebRTC is peer-to-peer, so it scales for free" confuses the API with the topology. A broadcast to 10,000 viewers cannot be a mesh in practice: the broadcaster's uplink could carry only a handful of copies. Every large WebRTC audience is served by media servers, and their cost per viewer is the number that decides the design.
- Why can a CDN cache an HLS segment for thousands of viewers but not WebRTC media?An HLS segment is a file fetched by an HTTP GET, identical for every viewer, so one cached copy serves them all. WebRTC media is pushed as RTP packets over UDP inside a stateful session, and because SRTP keys differ per connection, the bytes each viewer receives are different. There is no shared object to cache.
- In a hybrid design, how does a viewer watching over HLS get on stage to ask a question?The application connects that viewer to the WebRTC SFU hosting the stage, so their audio and video reach the speakers in under a second and they hear the reply without delay. The passive audience still sees the exchange through the composed HLS or DASH output, a few seconds later. When the viewer leaves the stage, they return to the HLS player.
saying these in an interview costs you the question
- WebRTC scales to any audience for free because it is peer-to-peer.
- HLS is slow because HTTP itself adds seconds to every request.
- A CDN can cache WebRTC packets the same way it caches segments.
- A WebRTC broadcast can switch encryption off to save server CPU.
- Every live stream needs sub-second latency, so WebRTC is always better.