For WebRTC meetings whose participants span three continents, would you cascade SFUs across regions or host each meeting on one SFU, and why?
answer
- short access legs matter most
- one copy per region, not per viewer
- loss recovery round trips
- nobody standardised SFU-to-SFU
- complexity only when measured need
basics
~20 sCascade when participants are spread far apart: each joins the nearest SFU over a short access path, and SFUs exchange one copy of each needed stream between regions. The price is an unstandardised SFU-to-SFU layer, extra hops and harder failure handling.
solid answer
~50 sWith one SFU in Europe, an Australian participant sends and receives everything over a long public path: retransmission and key-frame requests take a full long round trip, jitter buffers grow, and each Australian receiver pulls its own copy of every stream across the ocean. Cascading puts an SFU near each cluster: the access leg is short, loss is repaired quickly, and each stream crosses regions once per receiving SFU, which fans it out locally. The costs are real: the WebRTC specifications define only the participant-to-server legs, so the SFU-to-SFU protocol, room state and layer demand are each implementation's own; every extra hop adds latency and another SRTP decrypt, so without SFrame more servers can read media; and failover gets harder. I cascade large or genuinely global meetings and keep small regional ones on one nearby SFU.
go deeper
Know that a server far from a participant makes the call worse for them, and that several servers in different places can share one meeting.
Explain how a cascade works: nearest SFU per participant, one copy per stream between servers, local fan-out to each region's receivers.
Discuss loss recovery over short versus long round trips, layer demand between SFUs, and how a regional SFU failure is detected and survived.
Own the trade: a proprietary inter-server layer, room-state consistency and data-residency rules against better quality for distant users, and set the measured threshold for cascading.
## The problem with one SFU for a global meeting Picture a weekly meeting with eight people in Europe, five in North America and four in Australia, hosted on one SFU in Europe. Every Australian participant sends and receives all media over a long public-internet path: - **Loss recovery is slow.** A retransmission request or a key-frame request travels the whole long round trip before anything is repaired, so a lost packet freezes video for longer. - **Jitter is higher.** Long paths cross more networks, so receivers need bigger jitter buffers, adding delay to every stream. - **Cross-ocean traffic is duplicated.** Each of the four Australian receivers pulls its own copy of every European stream across the same long path. None of this changes the SFU's basic job of forwarding without decoding; it changes **where** that forwarding happens. ## What cascading does In a **cascaded** design, several SFUs serve one meeting: 1. Each participant connects to the **nearest SFU**, so the access leg — where most loss and jitter live — is short. 2. The SFUs connect to each other, often over better-provisioned inter-region links. 3. Each stream a remote region needs crosses between regions **once per receiving SFU**, and that SFU fans it out to its local participants. 4. A retransmission request for a packet lost on the access leg can be answered by the local SFU over a short round trip, if it keeps recent packets; key-frame requests still have to reach the sender. The WebRTC specifications define the participant-to-server legs (offer/answer semantics, ICE, DTLS-SRTP, simulcast negotiation). **How two SFUs exchange media, room membership and layer demand is not standardised**; each implementation designs it. That is the first thing a lead must own. ## What it costs | Concern | One SFU | Cascaded SFUs | |---|---|---| | Access path for distant users | long, lossy | short | | Cross-region copies per stream | one per remote receiver | one per remote SFU | | Extra media hops | none | one or more per remote stream | | Room state | in one process | distributed, must stay consistent | | Failure domain | one server ends the meeting | a region can fail alone, but needs re-routing | | Operational complexity | low | high: placement, monitoring, routing | The added hop costs a little latency and another decrypt and re-encrypt, since DTLS-SRTP protects one hop at a time. Without SFrame, every SFU in the chain can read the media, which widens the set of servers and jurisdictions that see content; with SFrame end-to-end encryption, none of them can, and cascading changes nothing for content confidentiality. ## Decisions a lead owns - **Placement.** Start the meeting on the region nearest most participants, and add a remote SFU only when a remote cluster is large or far enough to justify it. A meeting of three people in one city should never cascade. - **Layer demand.** A relaying SFU should pull from its upstream peer the highest simulcast encoding or SVC layer that any local receiver currently needs, and drop to lower ones when no one needs them; pulling every layer wastes inter-region capacity, pulling too few forces slow upgrades. - **Speaker selection.** Forwarding only active speakers and thumbnails across regions keeps the inter-region volume roughly flat as meetings grow. - **State and signalling.** Room membership, who publishes what, and key-frame requests must flow between SFUs reliably; a stale view produces black tiles. - **Failover.** Decide what happens when a regional SFU dies: participants reconnect to the next region, and the cascade re-forms. - **Data residency.** Some customers require media to stay in a jurisdiction; cascading can honour that by keeping their participants' SFU local, or violate it by relaying through a region they forbid. ## Measure before you cascade The decision should rest on data the platform already collects: - round-trip time from each participant to the SFU hosting the meeting; - loss and retransmission rates on those access legs; - how often meetings contain a remote cluster of more than one or two participants; - where customers' residency rules allow media to flow. ## When not to cascade Most meetings are small and regional. For them, a single SFU near the participants is simpler, cheaper and equally good. Cascading earns its complexity for large meetings, broadcasts with globally spread audiences, and products whose users are routinely split across continents — and it should be justified with measurements of where participants actually are, not adopted by default.
- In a cascade, which simulcast encoding should a relaying SFU pull from its upstream SFU?The highest encoding, or SVC layer, that any of its local receivers currently needs, and nothing above it. When the last receiver needing it downgrades, the relaying SFU drops to the lower one. Pulling every encoding wastes inter-region capacity; pulling only the lowest makes upgrades wait for a round trip and a key frame.
- Does cascading SFUs weaken the confidentiality of meeting media?With DTLS-SRTP alone, yes in scope: each SFU terminates a hop and holds plaintext encoded media, so more servers, possibly in more jurisdictions, can read it. With SFrame end-to-end encryption, no SFU can decrypt the media, so cascading adds no content exposure, though metadata such as who speaks and when stays visible to each SFU.
saying these in an interview costs you the question
- Cascading always lowers latency because more servers share the load.
- Each remote receiver should pull its own copy from the origin SFU.
- WebRTC standardises the protocol two SFUs use to cascade a meeting.
- A cascade requires MCU-style decoding where regions meet.
- Every meeting should be cascaded across all regions by default.