skip to content

WebRTC

Browser-native peer connections that carry encrypted media and data directly between peers once an SDP offer/answer and ICE checks succeed. Interviewers ask why a P2P call still needs servers.

part ofAPI stylesoverview, primer and where to startread it →
on this pageshow

explore

questions

26

A one-to-one WebRTC video call is pitched as 'serverless'; which server roles does it still need, and why does each one exist?

level: juniorimportance: must knowfreq 45%

answer

  1. only the media path is direct
  2. who carries the offer and answer
  3. learning your outside address
  4. relay when no pair connects
  5. groups, recording and telephony bridges

basics

~20 s

WebRTC still needs a signalling service to carry offers, answers and candidates, which the standards leave to the application; STUN to learn public addresses; TURN to relay when no direct path works; and media servers only for groups, recording or gateways.

solid answer

~50 s

Only the **media path** is peer-to-peer. Before it exists, the two browsers must exchange session descriptions and ICE candidates, and neither JSEP (RFC 9429) nor the W3C API says how: the application runs a **signalling** service, usually over HTTPS or WebSocket. A **STUN** server lets each peer learn the public address its NAT gives it, so it can advertise a candidate the far side can reach. A **TURN** server relays the still-encrypted packets when no direct pair succeeds; RFC 8835 makes TURN support mandatory, including TURN over TCP and over TLS for networks that block UDP. **Media servers** are optional for one-to-one: they appear for group calls, server-side recording or bridging to telephony, and each one joins the call as a WebRTC endpoint. The honest pitch is 'peer-to-peer media, server-assisted setup, relay as a fallback'.

go deeper

for a junior

Name the four roles and one job for each: signalling exchanges offers and candidates, STUN reveals your public address, TURN relays, media servers handle groups or recording.

for a middle

Explain that the standards leave signalling to the application, that STUN carries no media, and that a TURN relay forwards encrypted packets it cannot read.

for a senior

Show which servers sit on the critical path of every call, where the bandwidth cost lives, and why adding recording or group calls turns a media server into a party to the call.

for a principal

Frame the build-versus-operate choice: a self-run signalling and TURN estate against managed relays, and how relay share, regions and recording needs drive cost and privacy commitments.

## What "peer-to-peer" actually covers WebRTC is often described as browser-to-browser, and for the **media path** that is true: once a call is up, audio, video and data can flow directly between the two endpoints. RFC 8825, the WebRTC overview, draws this as a **trapezoid**: the media path ("low path") runs between the browsers, while the **signalling path** ("high path") runs through web servers that the application operates. So a "serverless" pitch is right about where the media usually goes and wrong about everything that has to happen before and around it. Four server roles exist, and each answers a different question: - **Signalling** answers "how do the two peers find each other and agree on a session?" - **STUN** answers "what address does the outside world see for me?" - **TURN** answers "how do we talk when no direct path works?" - **Media servers** answer "what if the call is more than two people talking?" ## Signalling: the role the standards leave to you Two browsers cannot open a connection to each other out of nowhere: neither knows the other's address, codecs, ICE credentials or DTLS certificate fingerprint. They must exchange **session descriptions** (an offer and an answer) and **ICE candidates** first. JSEP (RFC 9429) is explicit that it "does not specify a particular signaling model": it creates and applies descriptions, and how they travel — addressing, retransmission, forking, glare handling — is left to the application. The W3C WebRTC specification likewise says the signalling channel is "provided by unspecified means", typically a script talking to a server over WebSocket or HTTP. In practice the signalling service also carries everything around the call: who is online, who may call whom, ringing, and a hang-up message. It is always needed, even when both peers sit on the same desk. ## STUN and TURN: getting the packets through Most devices sit behind NATs and firewalls, so their own interface addresses are not reachable from outside. WebRTC uses ICE to gather candidate addresses and test pairs of them; two kinds of server feed it: 1. A **STUN** server answers a binding request with the address and port the request arrived from, so a peer learns its **server-reflexive** address and can offer it as a candidate. STUN carries no media. 2. A **TURN** server allocates a **relayed** address and forwards packets between the peers when no direct pair succeeds. RFC 8835 makes TURN support mandatory for WebRTC endpoints, and requires the TCP and TLS-over-TCP modes so a call can survive a network that blocks UDP. A relay costs bandwidth and adds latency, but it does not weaken encryption: the DTLS handshake runs between the two peers through it, so the relay forwards packets whose keys it never holds. RFC 8825 describes such relays as entities that "handle the data but do not modify it". ## Media servers: when the call stops being one-to-one A plain one-to-one call needs no media server. One appears when the product needs something two endpoints cannot do alone: - **group calls**, where forwarding or mixing streams centrally beats every participant uploading to every other; - **server-side recording** or transcription; - **gateways** to the telephone network, or to SIP equipment that lacks WebRTC's ICE and security mechanisms; - **broadcast** to large audiences. RFC 8825 notes that nothing restricts the protocols to browser-to-browser use: any endpoint that implements them faithfully interoperates. A media server is exactly that — it terminates each participant's peer connection as a WebRTC endpoint itself, which also means it is a party to the media rather than a pass-through. ## Who sees what | Server role | Needed when | Carries | Sees decrypted media? | |---|---|---|---| | Signalling | always | offers, answers, candidates, call control | no | | STUN | almost always (any NAT) | binding requests and responses | no media at all | | TURN | when no direct pair works | relayed DTLS and SRTP packets | no | | Media server | groups, recording, gateways | the media, as a call endpoint | yes | ## Taking the pitch apart A fair rewrite of "serverless video calling" is: **peer-to-peer media, server-assisted setup, relay as a fallback**. The operational consequences follow from the table: - the signalling service is on the critical path of every call, so it needs the availability of any login-gated API; - STUN is cheap, because it answers a few small requests per call; - TURN is where the bandwidth bill lives, and its share of calls depends on users' networks, not on the code; - adding recording or group calls changes the architecture, because a media server becomes a party to the call.

  • Can two WebRTC peers on the same office network connect with no STUN or TURN server configured?
    Usually yes. With an empty `iceServers` list the ICE agent still gathers **host** candidates, the device's own interface addresses, and two peers on one network can often reach each other on those. Signalling is still required to exchange descriptions and candidates. The setup breaks as soon as a NAT or a firewall sits between them, which is why deployments configure STUN and TURN.
  • Can the operator of a TURN relay listen to a relayed WebRTC call?
    Not from the relayed packets. The DTLS handshake runs between the two peers through the relay, so the relay forwards DTLS and SRTP packets whose keys it never holds. It does learn metadata: both peers' addresses, timing and traffic volume. A media server is different: it terminates each participant's connection, so it does see the media.
  • Why does WebRTC not standardise a signalling protocol?
    JSEP (RFC 9429) deliberately decouples creating and applying descriptions from how they travel, so an application can reuse its own HTTPS or WebSocket API, or bridge to SIP or XMPP as RFC 8825's trapezoid shows. The cost is that every deployment builds and secures its own rendezvous, presence and authorisation, and two separately built applications cannot call each other without agreeing on one.

Two pen-pals who know only each other's names ask a mutual friend to pass along their addresses (signalling); each first asks the post office what its return address looks like from outside (STUN); if one building refuses unknown mail, they use a forwarding box at the post office (TURN). Once letters flow, the friend steps out.

saying these in an interview costs you the question

  • WebRTC is fully peer-to-peer, so no server is involved at all
  • The STUN server relays the media when a direct path fails
  • Every browser speaks one standard WebRTC signalling protocol
  • A TURN relay decrypts the media, so it can read the call
  • Every WebRTC call, even one-to-one, has to pass through a media server
open as a page

In WebRTC, when would a browser game send player inputs over an RTCDataChannel rather than over a WebSocket to a server?

level: juniorimportance: must knowfreq 32%

basics

~20 s

Pick an RTCDataChannel when late inputs are worth dropping: it sends SCTP messages over DTLS and UDP, optionally unordered and partially reliable, often peer to peer. A WebSocket is one ordered, fully reliable TCP stream to a server, simpler to set up.

open as a page

In WebRTC, how do you list STUN and TURN servers in an RTCPeerConnection's iceServers, and why must a TURN entry carry a username and credential?

level: juniorimportance: must knowfreq 32%

basics

~20 s

Each RTCIceServer entry in RTCConfiguration.iceServers names a server by a stun:, turn: or turns: URI. A STUN entry needs nothing more; a TURN entry must also carry username and credential, because the TURN server authenticates every Allocate before it spends relay capacity.

open as a page

In WebRTC, what do an SDP offer and an SDP answer contain, and what must the application carry between the peers itself?

level: juniorimportance: must knowfreq 40%

basics

~20 s

An SDP offer lists the media and data sections a peer proposes, with codecs, directions, ICE credentials and a DTLS fingerprint; the answer accepts or rejects each one. WebRTC never transports them: the application's own signalling channel carries offer, answer and candidates.

open as a page

In WebRTC's RTCDataChannelInit, what do ordered, maxRetransmits and maxPacketLifeTime control, and why may only one limit be set?

level: middleimportance: must knowfreq 24%

basics

~20 s

The ordered option chooses send-order or arrival-order delivery; maxRetransmits caps how often a lost message is resent, and maxPacketLifeTime caps how many milliseconds it may be sent or resent. Neither gives a reliable channel; both throws a TypeError.

open as a page

In WebRTC's JSEP (RFC 9429), in what order do offerer and answerer call createOffer, setLocalDescription, setRemoteDescription and createAnswer, and why does it matter?

level: middleimportance: must knowfreq 30%

basics

~10 s

Offerer: createOffer, setLocalDescription, send. Answerer: setRemoteDescription, createAnswer, setLocalDescription, send back. Offerer: setRemoteDescription. Creating only drafts SDP; setting applies it and moves the signalling state, so steps out of order are rejected.

open as a page

In a 12-person WebRTC video meeting, how do mesh, SFU and MCU topologies differ in client bandwidth and server CPU?

level: middleimportance: must knowfreq 35%

basics

~20 s

A mesh makes each of 12 peers upload 11 streams; an SFU takes one upload per peer and forwards it undecoded to the others; an MCU decodes, mixes and re-encodes everything, so clients get one stream but the server pays heavy CPU.

open as a page

Can a WebRTC application switch encryption off to save CPU or let a network probe record calls, and what does RFC 8827 require instead?

level: middleimportance: should knowfreq 27%

basics

~20 s

No. RFC 8827 requires all WebRTC media to be sent as SRTP keyed with DTLS-SRTP, forbids plain RTP and NULL-encryption cipher suites, secures every data channel with DTLS, and forbids the API from exposing the keys to the page's script.

open as a page

In WebRTC, what does one RTCPeerConnection bundle under its API, and how are its transport objects layered on each other?

level: middleimportance: should knowfreq 22%

basics

~20 s

One RTCPeerConnection bundles an ICE agent (RTCIceTransport), a DTLS transport layered on it (RTCDtlsTransport), RTP transceivers whose senders and receivers send SRTP over that DTLS transport, and, when data channels exist, an SCTP association (RTCSctpTransport) running inside DTLS.

open as a page

In WebRTC, why does a camera feed travel as a media track rather than as encoded frames sent over an RTCDataChannel?

level: middleimportance: should knowfreq 16%

basics

~20 s

A media track travels as RTP, carrying timestamps for playout and lip sync, RTCP feedback such as NACK and Picture Loss Indication, and mandatory rate adaptation. A data channel moves opaque SCTP messages, leaving the application to rebuild all of that.

open as a page

In WebRTC, what starts ICE candidate gathering on an RTCPeerConnection, how does each candidate reach the application, and how is the end of gathering signalled?

level: middleimportance: should knowfreq 20%

basics

~20 s

Applying a local description with setLocalDescription starts a gathering phase. Each host, server-reflexive or relay candidate then arrives in an icecandidate event; an empty candidate string marks one transport's end-of-candidates, and iceGatheringState reaching complete marks the whole connection's.

open as a page

Why should a WebRTC application give browsers short-lived TURN credentials instead of one fixed username and password, and how is such a credential usually built?

level: middleimportance: should knowfreq 16%

basics

~20 s

Any TURN credential a browser receives is readable by whoever loads the page, so a fixed password turns the relay into free, identity-masking bandwidth for anyone. Credentials minted per user by the backend and expiring soon limit that abuse to a short window.

open as a page

In WebRTC, when a peer adds a screen-share track to a running call, what changes in the next SDP offer and what must stay fixed?

level: middleimportance: should knowfreq 18%

basics

~20 s

Adding the track fires negotiationneeded. The next offer keeps each existing m= section at its index and mid, adds a video section (appended or in a recycled port-0 slot), and puts its mid in the BUNDLE group, sharing the existing transport.

open as a page

With trickle ICE (RFC 8838) in WebRTC, how do candidates travel over the signalling channel, and what must the receiver do with ones arriving early?

level: middleimportance: should knowfreq 22%

basics

~20 s

The offer goes out at once; each candidate follows as its own message carrying its candidate line, sdpMid and ufrag. The receiver feeds them to addIceCandidate, queueing any that arrive before its remote description is set.

open as a page

In WebRTC, how is RTCPeerConnection's connectionState derived from its ICE and DTLS transports, and how should an application react to 'disconnected', 'failed' and 'closed'?

level: seniorimportance: should knowfreq 14%

basics

~20 s

connectionState aggregates ICE and DTLS: 'failed' if ICE or any DTLS transport failed, 'disconnected' if ICE is, 'connected' when both are up. Wait out 'disconnected', restart ICE on ICE failure, treat DTLS failure as a fault; 'closed' follows only your own close().

open as a page

Over WebRTC data channels, a browser game sends 2 MB map snapshots beside small inputs; why do the inputs stall, and how should the snapshots be sent?

level: seniorimportance: should knowfreq 11%

basics

~20 s

Base SCTP cannot interleave messages, so a 2 MB message's fragments monopolise the association and inputs on every channel wait; it may also exceed maxMessageSize and throw. Send snapshots in chunks of about 16 KB and pace them with bufferedAmount.

open as a page

When a WebRTC caller's laptop moves from office Wi-Fi to a phone hotspot mid-call, how does ICE detect the dead path, and how does an ICE restart recover the call?

level: seniorimportance: should knowfreq 14%

basics

~20 s

Consent freshness (RFC 7675) sends authenticated STUN checks on the selected pair; after 30 seconds without a response the endpoint must stop sending. An ICE restart, an offer with new ufrag and password, gathers candidates on the new network and selects a new pair.

open as a page

A WebRTC call fails for users on a corporate network that blocks outbound UDP and allows only TCP to port 443; how would you configure TURN, and what does that fallback cost?

level: seniorimportance: should knowfreq 12%

basics

~20 s

Add a turns: TURN entry on TCP port 443 beside the UDP and TCP ones. The browser reaches the relay over TLS, which such firewalls usually pass, while the relay speaks UDP to the peer; the price is head-of-line delay, TLS overhead and relayed bandwidth.

open as a page

In WebRTC, both peers send an offer at the same moment after each adds a screen share; what is glare, and how does perfect negotiation resolve it?

level: seniorimportance: should knowfreq 12%

basics

~20 s

Glare is receiving an offer while your own is still unanswered. Perfect negotiation, the W3C WebRTC pattern, pre-assigns roles: the impolite peer ignores the colliding offer, while the polite peer rolls its own offer back, answers, and negotiates its own change afterwards.

open as a page

In WebRTC, how does an SFU use simulcast or SVC to give each receiver a video quality that matches its bandwidth?

level: seniorimportance: should knowfreq 16%

basics

~20 s

With simulcast the sender encodes the same video at several resolutions as independent streams and the SFU forwards one to each receiver; with SVC it sends one layered stream and the SFU drops upper layers. Neither makes the SFU decode.

open as a page

When would you deliver a 10,000-viewer live session over WebRTC through SFUs rather than HLS or DASH, and at what cost?

level: seniorimportance: should knowfreq 20%

basics

~20 s

Choose WebRTC when viewers must react within a second, as in live bidding or on-camera questions. HLS or DASH take seconds but scale on cacheable HTTP; WebRTC needs a stateful, separately encrypted session per viewer on SFU capacity.

open as a page

In WebRTC, how does a data channel opened in-band with DCEP differ from one created with negotiated: true and an agreed id?

level: middleimportance: nice to knowfreq 9%

basics

~20 s

In-band, createDataChannel sends a DCEP DATA_CHANNEL_OPEN with the channel's options; the peer replies DATA_CHANNEL_ACK and its application gets a datachannel event. With negotiated: true, each application creates the channel itself on an agreed id, and no DCEP message is sent.

open as a page

In WHIP (RFC 9725), how does a WebRTC encoder publish a stream with one HTTP POST, and what does the 201 Created response carry?

level: middleimportance: nice to knowfreq 6%

basics

~20 s

The encoder POSTs its SDP offer as application/sdp to the WHIP endpoint URL. The endpoint replies 201 Created with the SDP answer as the body and a Location header naming the session resource, later used for PATCH (ICE updates) and DELETE (teardown).

open as a page

In WebRTC, what does setting iceTransportPolicy to "relay" change about ICE candidates, and when is forcing every call through TURN worth its cost?

level: seniorimportance: nice to knowfreq 7%

basics

~20 s

With iceTransportPolicy relay, the ICE agent surfaces and checks only relay candidates, so the far peer never learns this endpoint's own addresses. It costs relay bandwidth and a detour on every call, and the call fails outright if no TURN server answers.

open as a page

In a WebRTC call through an SFU, why is DTLS-SRTP not end-to-end encryption, and how does SFrame (RFC 9605) add it?

level: seniorimportance: nice to knowfreq 7%

basics

~20 s

DTLS-SRTP keys each hop separately, so an SFU decrypts every packet and can read the media. SFrame encrypts each encoded frame with keys the SFU never holds, inside the hop-by-hop SRTP, leaving the SFU only the metadata it needs to forward.

open as a page

For WebRTC meetings whose participants span three continents, would you cascade SFUs across regions or host each meeting on one SFU, and why?

level: principalimportance: nice to knowfreq 9%

basics

~20 s

Cascade when participants are spread far apart: each joins the nearest SFU over a short access path, and SFUs exchange one copy of each needed stream between regions. The price is an unstandardised SFU-to-SFU layer, extra hops and harder failure handling.

open as a page