What breaks if unmodified TLS records are sent over UDP, and what does DTLS change to fix it?
answer
- nothing underneath repairs anything
- records must stand alone
- the counter moves onto the wire
- the handshake gets its own timer
- epoch, message_seq, fragment_offset
basics
~20 sTLS assumes an ordered, reliable byte stream: drop or reorder one record and decryption fails, and a lost handshake message stalls forever. DTLS puts that back into the protocol itself with explicit per-record numbering, fragmented handshake messages and its own retransmission timer.
solid answer
~50 sThe TLS record layer was written on top of a reliable ordered stream, and that assumption is load-bearing. The per-record counter that feeds the AEAD nonce is never transmitted — both ends simply count — and a handshake message is just a run of bytes in a stream that always arrives. Point that at a datagram socket and one dropped or reordered datagram desynchronises the counter, so every later record fails its integrity check with `bad_record_mac(20)`; a lost handshake message stalls the handshake permanently because nothing underneath retransmits. DTLS keeps the same message order, extensions and ciphers and replaces only the transport assumptions: an `epoch` and an explicit `sequence_number` in every record so each record decrypts on its own, `message_seq` / `fragment_offset` / `fragment_length` so one handshake message can span datagrams, a retransmission timer per flight, and a stateless cookie exchange because a datagram source address is unverified.
code
pseudocode · 12 linesTLS record header over a stream (5 bytes)
type 1 byte
version 2 bytes
length 2 bytes
-- the record sequence number is counted by both ends and never sent
DTLS 1.2 record header over datagrams (13 bytes)
type 1 byte
version 2 bytes
epoch 2 bytes which key set protects this record
sequence_number 6 bytes counted per epoch, sent on the wire
length 2 bytesgo deeper
Remember the one-line reason DTLS exists: TLS assumes something beneath it already fixed loss and reordering, and over datagrams nothing did. Name at least the explicit sequence number and the handshake retransmission timer.
Explain the mechanics: the nonce is built from a counter both ends infer, so a gap desynchronises it and every later record fails integrity. Then list what DTLS adds to the record header and to the handshake header.
Show where the boundary sits in production. Only handshake flights are retransmitted, so application loss is still the application's problem, and that is a feature for loss-tolerant senders rather than a gap to be papered over.
Frame it as a transport choice. Picking DTLS means accepting unrepaired loss in exchange for no head-of-line blocking and no connection setup cost per burst, and it commits the fleet to a handshake whose cost and timers you must plan for on constrained links.
## The assumption TLS is built on TLS protects a byte stream. Everything beneath it is assumed to have already solved delivery: every record arrives, arrives exactly once, and arrives in the order it was sent. Three design choices follow from that single assumption, and all three are what a buoy telemetry uplink over a satellite or cellular link immediately violates. - **The record sequence number is never sent.** Both peers keep a 64-bit counter and increment it per record. It is mixed into the per-record AEAD nonce, so sender and receiver only produce the same nonce while their counters agree. - **There is no retransmission and no acknowledgement.** A TLS implementation that has sent a `Finished` message simply waits; if it never arrives, that is somebody else's problem. - **A handshake message is not bounded by a record.** A large `Certificate` message can be split across records arbitrarily, because the stream reassembles them for free. ## What actually happens on a datagram socket Run those rules over datagrams and each one fails in a different way: 1. **One dropped record desynchronises the counter.** The receiver builds its nonce from a number the sender has already moved past, so authenticated decryption fails — and keeps failing for every subsequent record. The symptom is not 'one lost reading' but a dead association. 2. **Reordering looks identical to corruption.** Datagrams routinely arrive out of order; with an inferred counter there is no way to say *which* record this is. 3. **A lost handshake message is fatal and silent.** Neither side retransmits, so both wait: the client for a `ServerHello`, the server for a `Finished` that will never come. 4. **An oversized handshake message cannot be sent at all.** A datagram must fit the path MTU. A certificate chain of two or three kilobytes has nowhere to go. ## What DTLS puts back DTLS is deliberately not a new security protocol. It reuses the TLS message order, the extension set, the key exchange and the ciphers, and changes only what the missing transport guarantees used to provide: 1. **Explicit record numbering.** Every record carries an `epoch` (which key set protects it) and a `sequence_number` on the wire, so each record is an independently decryptable unit. 2. **A handshake reliability layer.** The handshake header gains `message_seq`, `fragment_offset` and `fragment_length`, so one message can be split across datagrams and reassembled, and messages can be ordered and queued when they arrive out of order. 3. **A retransmission timer per flight.** DTLS groups handshake messages into flights and retransmits a whole flight when its timer expires. 4. **Duplicate detection.** A sliding receive window lets a receiver discard a record it has already accepted, silently. 5. **A stateless cookie exchange.** Because a datagram source address is trivially spoofed, the server proves return routability before allocating any state. ## What DTLS deliberately does not do This is where most wrong answers live. DTLS does **not** turn the datagram service into a reliable ordered one: - **Application records are still lost, still reordered and still unrepaired.** Only handshake flights are retransmitted. A telemetry burst that is dropped is gone, and that is the property a buoy or a real-time media sender wants — a stale reading retransmitted ten seconds later is worth less than the next one. - **It does not reorder for you.** A receiver sees records in arrival order with their numbers attached; deciding what to do about a gap is the application's job. - **Ciphers whose keystream depends on every preceding byte cannot be used**, because a receiver cannot skip a gap and resynchronise. ## Side by side | Property | TLS over a stream | DTLS over datagrams | |---|---|---| | Record sequence number | counted by both ends, never sent | carried in every record with an `epoch` | | Record header size | 5 bytes | 13 bytes in DTLS 1.2 | | Lost handshake message | cannot happen | retransmitted on the sender's timer | | Handshake message larger than one packet | split across records by the stream | split with `fragment_offset` / `fragment_length` | | Lost application data | repaired beneath TLS | not repaired at all | | Duplicate record | cannot happen | discarded by the replay window | The compact way to hold it: **TLS delegates delivery downwards; DTLS has nobody to delegate to, so it carries the minimum it needs to survive loss, reordering and duplication itself — and refuses to pretend it has fixed them for your data.**
- Does DTLS retransmit a lost application record the way it retransmits a handshake flight?No. Retransmission in DTLS covers handshake flights only. Application records are delivered with exactly the reliability the datagram service gives them: they can be lost, duplicated or reordered, and DTLS reports none of it. An application that needs reliability has to build it above DTLS, which is precisely why loss-tolerant senders such as telemetry bursts and real-time media choose this transport.
- If DTLS changes so little of TLS, why is it a separate specification rather than an option?Because the changes are in the record and handshake headers, which every implementation parses before anything else. An `epoch`, an explicit sequence number and the three fragmentation fields change the wire format of every record and every handshake message, and the flight-based retransmission state machine has no counterpart in TLS. It is the same cryptography with a different framing and a different state machine.
- Why does a dropped record in stream-mode TLS never cause this problem in the first place?Because the loss is repaired below TLS. The transport retransmits the missing bytes and hands the record layer a gap-free stream, so the implicit counters on both sides stay in step and the record layer never observes that anything went wrong. Remove that repair and the counters are the first thing to break.
A stream protocol is dictation over a phone line; a datagram protocol is postcards. Postcards need a number written on each one, and the sender has to notice when a reply never comes.
saying these in an interview costs you the question
- Says DTLS makes UDP reliable and ordered, like TCP
- Claims a lost telemetry record is retransmitted by DTLS
- Describes DTLS as TLS with the socket type swapped, no wire change
- Thinks DTLS uses different cryptography or a different message order
- Cannot say why an inferred sequence counter fails over datagrams
- Treats DTLS and QUIC as the same UDP security protocol