In the 12-byte RTP fixed header, what do the payload type, sequence number, timestamp and SSRC each tell a receiver?
answer
- codec, order, sampling clock, source
- sequence +1 per packet sent
- timestamp counts samples, not milliseconds
- 160 ticks per 20 ms at 8 kHz
- SSRC random, CNAME binds it
basics
~20 sThe payload type names the codec, the sequence number (+1 per packet) reveals loss and reordering, the timestamp counts media-clock ticks of the sampling instant for playout and jitter, and the SSRC identifies which source the packet belongs to.
solid answer
~50 sThe 7-bit `payload type` says how to decode the payload — a static number such as 0 for PCMU (RFC 3551) or a dynamic 96-127 value bound in SDP; a receiver MUST ignore types it does not understand. The 16-bit `sequence number` rises by one per packet sent, so gaps show loss and inversions show reordering. The 32-bit `timestamp` is the sampling instant of the first octet in units of the media clock — 8,000 Hz for PCMU, so 20 ms packets step by 160 — and it keeps advancing through suppressed silence. The 32-bit `SSRC` is a random identifier for the source; it is the key a receiver files the stream under. Sequence number and timestamp start at random values, and the timestamp is not wall-clock time: RTCP sender reports pair it with an NTP timestamp for synchronisation.
go deeper
Name the four fields and the one job of each: codec, packet order, sampling time, source identity.
Work the arithmetic: 160 ticks per 20 ms PCMU packet, a sequence wrap in about 22 minutes at 50 packets per second, and why silence moves the timestamp but not the sequence number.
Use the fields to diagnose: tell loss from silence suppression, spot an SSRC change after a transport move, and know which fields a forger must match and which SRTP needs in clear.
Weigh what the clear header exposes against what monitoring and header compression gain from it, and when that metadata exposure is acceptable.
## The layout RFC 3550 §5.1 defines a header whose first twelve octets are present in every RTP packet: | Bits | Field | Meaning | |---|---|---| | 2 | `V` | Version, always 2 | | 1 | `P` | Padding octets at the end of the packet | | 1 | `X` | One header extension follows the fixed header | | 4 | `CC` | Number of CSRC identifiers after the fixed header (0-15) | | 1 | `M` | Marker, meaning set by the profile | | 7 | `PT` | Payload type | | 16 | sequence number | +1 per RTP data packet sent | | 32 | timestamp | Sampling instant of the first payload octet | | 32 | `SSRC` | Synchronization source identifier | A mixer may append up to 15 32-bit **CSRC** identifiers naming the sources it mixed together. ## Payload type — how to decode The **payload type** identifies the payload format. RFC 3551, the audio/video profile, gives static numbers: 0 is PCMU, 8 is PCMA, both at an 8,000 Hz clock. The range **96-127** is for dynamic assignment, bound per session by signalling, so "96" means whatever this call's SDP said it means. A receiver MUST ignore packets with a payload type it does not understand, and RFC 3550 says the field SHOULD NOT be used to multiplex separate media streams in one session. ## Sequence number — order and loss The **sequence number** increments by one for each RTP data packet sent, so a receiver can detect loss (a gap) and restore order (an inversion). Its initial value SHOULD be random. At 16 bits it wraps quickly: a voice stream of 50 packets per second (20 ms each) wraps after 65,536 / 50 = **1,310.72 seconds**, about 21.8 minutes, so receivers count wrap cycles to keep an extended sequence number. ## Timestamp — the media clock The **timestamp** is the sampling instant of the first octet, counted in ticks of a clock whose rate the payload format fixes. Worked example for PCMU: 1. Clock rate 8,000 Hz; one packet carries 20 ms. 2. 8,000 x 0.020 = **160** ticks per packet. 3. Consecutive packets: sequence +1, timestamp +160. 4. If one 20 ms block is suppressed as silence, the next packet has sequence +1 but timestamp +320: RFC 3550 advances the timestamp "regardless of whether the block is transmitted in a packet or dropped as silent". So the sequence number reveals **loss**, the timestamp reveals **time**, and comparing the two separates loss from silence. A 32-bit timestamp at 8,000 Hz wraps after 2^32 / 8,000 = 536,870.9 seconds, about 6.2 days; RFC 3551's video encodings use a 90,000 Hz clock and wrap in about 13.3 hours. The initial value SHOULD be random, so the first wrap can come much sooner. The timestamp is **not wall-clock time**, and two streams' timestamps advance at different rates from random offsets. RTCP sender reports pair an RTP timestamp with an NTP-format wallclock timestamp, which is how a receiver lines up audio and video. ## SSRC — whose packet this is The **SSRC** is a random 32-bit identifier, chosen so that no two sources in one session collide. It defines a single timing and sequence-number space: a receiver keeps per-SSRC state for loss, jitter and playout. It is not stable for life — it changes on a collision, and a source that changes its transport address must choose a new one — so RTCP's SDES **CNAME** item binds the SSRC to a persistent endpoint identifier. ## The remaining bits - **Marker (`M`)**: for audio, RFC 3551 says it SHOULD be set on the first packet of a talkspurt after silence, and applications without silence suppression MUST leave it zero; RFC 3550 intends it for significant events such as video frame boundaries. - **`X` and `CC`**: announce a header extension and the number of CSRCs, so the header can grow beyond twelve octets. - **Reserved payload types**: RFC 3550 requires the octet holding `M` and `PT` to avoid the values 200 and 201, so payload types 72 and 73 are reserved, and RFC 3551 reserves 72-76, keeping RTP distinguishable from RTCP. ## Why this matters for security - Every one of these fields travels in clear in plain RTP, and SRTP's pre-defined transforms still leave them readable. - The SSRC and sequence number are what a forged packet must match to be accepted. - The payload type tells an eavesdropper which decoder to use. - SRTP builds its packet index from the sequence number, which is why the field cannot simply be encrypted.
- On a PCMU call with silence suppression, two consecutive sequence numbers carry timestamps 480 apart; what happened?Two 20 ms blocks were suppressed as silence, not lost. Each block is 160 ticks at 8,000 Hz; the gap is the previous packet's own block (160) plus the two unsent blocks (2 x 160), so 480, while the sequence number, which counts only packets sent, rose by one. RFC 3551 also says the first packet of a talkspurt SHOULD carry the marker bit.
- Why do RTCP sender reports exist if every RTP packet already carries a timestamp?RTP timestamps have random offsets and stream-specific clock rates, so audio and video timestamps cannot be compared directly. A sender report pairs one RTP timestamp with an NTP-format wallclock time, giving the receiver the mapping it needs to synchronise streams, plus packet and octet counts for statistics.
saying these in an interview costs you the question
- The RTP timestamp is the sender's wall-clock time in milliseconds.
- The sequence number and timestamp both start at zero on every call.
- The SSRC is derived from the sender's IP address and port.
- Payload type 96 always means the same codec on every call.
- A jump in timestamps between consecutive packets always means packets were lost.