In WebRTC, what do an SDP offer and an SDP answer contain, and what must the application carry between the peers itself?
answer
- two text blobs, one each way
- one m= section per transceiver
- codecs, direction, ICE details, fingerprint
- created and applied, never transported
- your channel, your messages
basics
~20 sAn SDP offer lists the media and data sections a peer proposes, with codecs, directions, ICE credentials and a DTLS fingerprint; the answer accepts or rejects each one. WebRTC never transports them: the application's own signalling channel carries offer, answer and candidates.
solid answer
~40 sBoth are Session Description Protocol text. The offer has one `m=` section per media transceiver, plus one `m=application` section if data channels are used; each carries its codecs (`a=rtpmap`), a direction (`a=sendrecv`, `a=sendonly`, ...), an `a=mid` identifier, ICE credentials (`a=ice-ufrag`, `a=ice-pwd`), any candidates gathered so far and a certificate fingerprint. The answer must contain exactly as many `m=` sections in the same order, and for each one either picks a compatible configuration or rejects it with port 0. JSEP (RFC 9429) specifies how descriptions are created and applied but not how they travel: addressing, retransmission, forking and glare are left to the application. So you choose the channel (a WebSocket, HTTP requests, a SIP or XMPP gateway) and send the offer, the answer and every trickled candidate over it, alongside your own call-control messages.
go deeper
Recall that the offer and answer are SDP text, one m= section per transceiver plus one for data, and that WebRTC leaves delivering them to the application's own channel.
Explain what an answer may change (codec choice, direction, rejection with port 0) and what it may not (the number and order of m= sections), and list what else travels on the channel.
Show you treat signalling as part of the system: authenticated, reliable, ordered delivery of descriptions and candidates, kept open for every later renegotiation and ICE restart.
Weigh what owning signalling buys, such as fitting SIP gateways, HTTP ingest or custom call control, against the protocol work JSEP hands you: addressing, retransmission, forking and glare.
## What the two descriptions are A WebRTC call is set up by exchanging two **session descriptions** written in the **Session Description Protocol (SDP)**, a line-oriented text format. The peer that starts a negotiation produces an **offer**; the other peer replies with an **answer**. This is the offer/answer model of RFC 3264, and the browser-side rules for producing and applying the two are **JSEP**, the JavaScript Session Establishment Protocol (RFC 9429, which obsoletes RFC 8829). Neither description carries media. Each one is a *proposal about* media: which streams exist, how they are encoded, which direction they flow, and how to reach and authenticate the peer that will send them. ## Anatomy of an offer An offer opens with session-level lines (`v=`, `o=`, `s=`, `t=`, and usually `a=group:BUNDLE`) and then holds one **`m=` section** (a media description) per transceiver the application has added, in the order they were added, plus one `m=application` section for data channels if any exist. | Line | What it tells the other peer | |---|---| | `m=audio ...` / `m=video ...` / `m=application ...` | the kind of stream and its transport profile | | `a=mid:0` | a stable identifier for this section | | `a=rtpmap:111 opus/48000/2` | a codec the sender is willing to receive, with its payload type | | `a=sendrecv`, `a=sendonly`, `a=recvonly`, `a=inactive` | the proposed direction of media | | `a=ice-ufrag`, `a=ice-pwd` | credentials for this side's ICE checks | | `a=candidate:...` | network addresses gathered so far (possibly none) | | `a=fingerprint:...` | the hash of the certificate the DTLS handshake must present | A codec list in a description says what that side is prepared to **receive**; intersected with the other side's list, it decides what each side sends. Some values are declarative rather than negotiated: the fingerprint is computed from the local certificate and is simply accepted or rejected. ## What the answer may and may not do The answer is constrained by the offer: - It **must contain exactly as many `m=` sections** as the offer, in the same order, so each answer section can be matched to the offer section at the same index. - For each section it chooses a configuration compatible with what was offered, such as a codec subset or a direction that fits (an offered `sendonly` can only be answered `recvonly` or `inactive`). - It may **reject** a section by giving it port 0; the section stays in place, and it is the port that marks it as rejected. - It **cannot add** a stream the offer did not describe. If the answerer wants to send something new, it makes its own offer in a later exchange. A description of type `pranswer` (a provisional answer) also exists, but most applications send a final answer directly. ## What JSEP leaves to you: the signalling channel RFC 9429 is explicit that the JSEP implementation is "totally decoupled" from how offers and answers reach the other side, "including addressing, retransmission, forking, and glare handling." The application decides: 1. **The transport.** A WebSocket to a signalling server is common, but HTTP requests, a SIP or XMPP gateway, or anything else that moves text will do. WHIP (RFC 9725) is one case where the transport is fixed: a single HTTP POST. 2. **The addressing.** Who is calling whom, and how the server routes a message to the right peer. 3. **What else travels.** Besides the offer and answer, each ICE candidate gathered after the offer was sent (trickle ICE) and an end-of-candidates indication, plus the application's own messages: ringing, accept, hang-up, role assignment. 4. **Reliability and ordering.** Trickle ICE requires each candidate to be delivered exactly once and in order, so the channel or the application's own protocol on top of it must provide that. The channel is not only for call setup. Adding a track, stopping one or restarting ICE all need another offer/answer exchange over the same channel. ## Why the design is split this way RFC 9429 records that a lightweight signalling protocol built into the API was considered and not chosen, because the implementation would then have to understand and handle concepts such as glare itself. Handing the application SDP instead keeps the implementation's job to generating and applying descriptions; the application converts them into the messages of whatever signalling protocol it uses, SIP and XMPP included, and back again. ## Common misreadings - Treating SDP as the media path. It describes media; RTP and SCTP packets flow separately once ICE and DTLS complete. - Assuming STUN or TURN servers deliver offers. They help ICE find and relay a network path; they never see the SDP. - Expecting the answer to mirror the offer word for word. It narrows codec lists, sets directions and may reject sections. - Discarding the channel after the call connects. Every later renegotiation and ICE restart needs it.
- Can a WebRTC answer add a media section that the offer did not contain?No. RFC 3264 and JSEP require the answer to hold exactly the same number of `m=` sections as the offer, in the same order, so each can be matched by index. If the answerer wants to send a new stream, it waits for the current exchange to finish and then makes its own offer containing the new section.
- Does the WebRTC signalling channel have to be a WebSocket?No. JSEP mentions WebSockets only as an example of a preferred signalling mechanism and leaves the choice to the application. Plain HTTP requests, a SIP or XMPP gateway, or a fixed protocol such as WHIP's single HTTP POST all work, provided they move the descriptions and trickled candidates reliably and in order.
- How does a WebRTC answerer turn down a video section it cannot handle?It keeps the section in its answer but sets the port on its `m=` line to 0, which marks that stream as rejected while preserving the section's position. The offerer sees the rejection when it applies the answer, and the associated transceiver is stopped. Later offers keep that slot, at port 0, until it is recycled for a new stream.
saying these in an interview costs you the question
- WebRTC defines its own signalling protocol and signalling server.
- The STUN server relays the offer and answer between peers.
- The SDP offer carries the encoded audio and video itself.
- An answer can add extra m= sections the offer never had.
- Signalling is needed only once, before the first media flows.