In a 12-person WebRTC video meeting, how do mesh, SFU and MCU topologies differ in client bandwidth and server CPU?
answer
- who sends how many copies
- n minus one uploads per peer
- forward versus decode, mix, encode
- one composite stream down
basics
~20 sA mesh makes each of 12 peers upload 11 streams; an SFU takes one upload per peer and forwards it undecoded to the others; an MCU decodes, mixes and re-encodes everything, so clients get one stream but the server pays heavy CPU.
solid answer
~40 sIn a **mesh**, every participant opens a separate `RTCPeerConnection` to every other one, so with 12 people each client uploads 11 encoded, separately encrypted copies of its media and decodes 11 incoming streams; uplink breaks first. An **SFU** (selective forwarding unit) receives one upload per participant and forwards the encoded packets to each receiver without decoding them, so client uplink stays flat and the server's cost is mostly egress bandwidth, up to 12 × 11 = 132 forwarded streams. An **MCU** decodes all 12 streams, composes one picture and audio mix, and re-encodes it, so each client sends one stream and receives one, which suits weak devices, but the server pays decode and encode CPU and adds delay. The SFU is the usual middle ground: bandwidth on the server, little CPU.
go deeper
Recall the three shapes: everyone sends to everyone, a server that forwards, a server that mixes. Be able to say why a mesh stops working as the room grows.
Do the arithmetic aloud: n − 1 uploads per peer in a mesh, one upload and up to n − 1 downloads with an SFU, one each way with an MCU. Explain forwarding without decoding.
Name where each design spends money, uplink, egress or CPU, and the hidden SFU costs: per-hop re-encryption, header rewriting, and no transcoding for a receiver lacking the sender's codec.
Frame the choice as a portfolio: mesh for tiny calls, SFU for interactive rooms, server audio mixing or an MCU for weak clients and legacy bridges, and say what each does to operating cost.
## Three ways to connect a group call A **topology** is the shape of the media paths in a call with more than two people. WebRTC's RTP usage document (RFC 8834, §5.1) names the options RTP itself allows: every endpoint sends to an **RTP middlebox** that redistributes the traffic, endpoints form a **mesh** of unicast streams, or they use IP multicast. It then narrows them for WebRTC: endpoints are not expected to support IP multicast or a mesh run as one RTP session, but a mesh built from **independent `RTCPeerConnection` objects**, one per remote peer, must be supported. The middlebox side splits into the two server designs interviewers ask about: - **SFU (selective forwarding unit)**: receives each participant's encoded media and forwards a selection of it to each receiver, without decoding. - **MCU (multipoint control unit)**: decodes the incoming media, mixes or composes it, and re-encodes the result. ## Mesh: the uplink grows with the room In a mesh of *n* participants each client keeps *n − 1* peer connections. Each connection has its own DTLS-SRTP keys and its own congestion control, so the client sends **n − 1 encoded, separately encrypted copies** of its camera and microphone and receives *n − 1* streams to decode. With 12 people and an illustrative 1.5 Mb/s per video stream, every participant must upload about **16.5 Mb/s** — more than most home and mobile uplinks sustain. Download grows the same way, and decode CPU grows with it. What a mesh saves is the server: apart from signalling and, when a direct path fails, a TURN relay that only forwards still-encrypted packets, no server decrypts or processes the media, so nothing in the middle can read it. That is why a mesh is still a sound choice for two, three or four people. ## SFU: forwarding without decoding An SFU turns the uplink problem into a server bandwidth problem. Each client uploads **once** (or a few simulcast encodings), and the SFU sends each receiver the streams it wants. For 12 participants that is up to **12 × 11 = 132 forwarded streams** of egress, though an SFU can forward fewer, for example only the active speakers. The SFU does real per-packet work but no per-pixel work: - it decrypts each packet with the sender leg's SRTP keys and re-encrypts it with each receiver leg's keys, because DTLS-SRTP protects one hop at a time; - it rewrites RTP header fields such as sequence numbers so each receiver sees one continuous stream when it switches sources (RFC 8853, §6.2); - it handles feedback such as key-frame requests. RFC 8853 (§3) sums up this "switching" approach: computationally cheap for the middlebox, little impact on quality of experience, no need for full access to the media content — at the price of a less perfect fit for each receiver and more upload from senders that offer several encodings. Because it never decodes, an SFU cannot transcode: every receiver must support the codec the sender negotiated. ## MCU: decode, mix, re-encode An MCU gives each client the simplest life — **one stream up, one stream down** — which suits weak devices and gateways to older systems. The server pays for it: it decodes all 12 inputs, composes a layout and mixes audio, then encodes the output, often per participant, because a listener should not hear their own voice back. RFC 8853 (§3) lists the costs of this "transcoding" approach: very computationally expensive, added end-to-end delay, and the server must access the media content. RFC 8834 (§5.1) also says video-switching MCUs **SHOULD NOT** be used in WebRTC, because they make RTCP congestion control problematic, while content-modifying MCUs that terminate RTCP **MAY** be used. ## The trade in one table | | Mesh | SFU | MCU | |---|---|---|---| | Uploads per client (12 people) | 11 | 1, or a few simulcast encodings | 1 | | Downloads per client | 11 | up to 11 | 1 | | Server media work | none | forward, rewrite headers, re-encrypt | decode 12, compose, encode | | Server sees decoded media | no | no (it does decrypt each hop) | yes | | Main cost | client uplink and CPU | server egress bandwidth | server CPU and added delay | ## Choosing 1. Two to four people on decent links: a mesh is cheapest and needs no media server. 2. Larger interactive meetings: an SFU, because its cost is bandwidth, which scales more cheaply than decode and encode CPU. 3. Very weak clients, a fixed composed layout, or bridging to a legacy endpoint: an MCU, or a hybrid in which an SFU forwards video and a server mixes only the audio.
- If an SFU never decodes video, why does it still do work on every packet?DTLS-SRTP protects one hop at a time, so the SFU decrypts each packet with the sender leg's keys and re-encrypts it for each receiver leg. It also rewrites RTP header fields such as sequence numbers so a receiver sees one continuous stream when the SFU switches sources, and it handles feedback such as key-frame requests. The work is per packet, not per pixel, which is why it is far cheaper than an MCU's decode and encode.
- When does a mesh still beat an SFU for a WebRTC call?For two to four participants with reasonable uplinks. There is no media server to run or pay for, the path is direct (or through a TURN relay when it must be), and no server ever decrypts the media because each peer connection's DTLS-SRTP runs between the two endpoints. Beyond that size, the n − 1 uploads per client exhaust typical uplinks and encryption and decode CPU.
An SFU is a news desk that routes each reporter's finished tape to the stations that want it, picking the long or short cut but never re-editing; an MCU is a studio that watches every tape and edits them into one programme.
saying these in an interview costs you the question
- An SFU decodes each stream and re-encodes it smaller for slow receivers.
- A mesh scales well because each peer downloads the same as with an SFU.
- An MCU is the cheapest server because it sends only one stream per client.
- Media through an SFU is end-to-end encrypted because the SFU never decodes it.
- A mesh call needs no servers at all, not even for signalling.