skip to content

A TLS 1.3 connection carrying one continuous stream stays open for hours - how are its record-protection keys rotated without a new handshake?

level: seniorimportance: should knowfreq 36%

answer

  1. rekey without repeating the handshake
  2. a protected message sent mid-stream
  3. each direction rotates on its own
  4. update_requested pulls the peer along
  5. sent under the old key, switch after

basics

~20 s

With a KeyUpdate message, sent inside the protected stream after the handshake. It rotates only the sender's own write keys and resets that direction's record counter; its request_update field, set to update_requested, asks the peer to rotate its direction too.

solid answer

~40 s

`KeyUpdate` is a post-handshake message carried in a protected record. The new secret for that direction is derived from the current application traffic secret with `HKDF-Expand-Label`, a fresh key and static IV follow from it, and the record counter for that direction resets to zero. The rotation is **one-directional**: sending `KeyUpdate` replaces the sender's write keys and nothing else. Its single field decides whether the peer is pulled along - `update_not_requested` rotates only this direction, while `update_requested` obliges the receiver to send its own `KeyUpdate` carrying `update_not_requested` before its next application data record, which is what stops the two sides from asking each other forever. The message itself is protected under the **current** key; the sender switches only after it has sent it.

code

pseudocode · 14 lines
pseudocode
function on_key_update(message):
    if handshake has not completed:
        reject("KeyUpdate before the handshake finished")

    read_secret = derive_next(read_secret)
    install_read_keys(read_secret)
    read_sequence_number = 0

    if message.request_update is update_requested:
        send_under_current_key(KeyUpdate with update_not_requested)
        write_secret = derive_next(write_secret)
        install_write_keys(write_secret)
        write_sequence_number = 0
    return

go deeper

for a junior

Know that TLS 1.3 can change record-protection keys on an open connection with a single message, rather than by starting the handshake again.

for a middle

Explain the ordering: the message goes out under the old key, the sender switches afterwards, the counter resets to zero, and only that one direction changes.

for a senior

Demonstrate operating it: rotate on a record or byte budget, keep accepting the peer's current key until its own update arrives, and recognise periodic integrity failures as an early-switch bug.

for a principal

The call to own is the rotation policy across a fleet of long-lived connections - what threshold, which direction, and whether a connection that cannot rotate in time is closed or allowed to run to its limit.

## Why a long-lived connection needs this at all A connection that carries a continuous feed for hours protects an enormous number of records under one key. Two separate limits make that a problem: - the 64-bit record counter for a direction must never wrap, and the first record under a new key starts it again at zero; - an authenticated-encryption algorithm may carry its own limit on how many records, or how many bytes, may safely be protected under a single key. Before TLS 1.3 the only way to install fresh traffic keys on an open connection was to run another handshake over it. TLS 1.3 removed that mechanism and replaced it with a single message. ## The message and its one field `KeyUpdate` is a post-handshake handshake message - its inner content type is `handshake(22)` - sent only after the handshake has completed. It carries exactly one field, `request_update`, whose two values are `update_not_requested` and `update_requested`. | value | what the sender is doing | what the receiver must do | |---|---|---| | `update_not_requested` | rotating its own write keys | rotate its read keys for that direction, and nothing more | | `update_requested` | rotating its own write keys and asking the peer to rotate the other direction | rotate its read keys, then send its own `KeyUpdate` with `update_not_requested` before its next application data record | The fixed reply value is the interesting design detail. If answering `update_requested` with another `update_requested` were permitted, two implementations that each always asked would trade updates forever; pinning the answer to `update_not_requested` ends the exchange after one round. ## Who switches what, and when 1. The sender protects the `KeyUpdate` message **under the key currently in force**, and only then installs its new write key and resets its write counter to zero. 2. The receiver opens that record under the key it already had, then installs the matching read key and resets its read counter to zero. 3. Nothing about the *other* direction has changed. The peer's records still arrive under the peer's existing key until it sends a `KeyUpdate` of its own. That third point is where implementations go wrong. After sending `update_requested`, an endpoint may receive any number of records from the peer before the peer's `KeyUpdate` turns up, because those records were already in flight when the request was sent. A receiver that tears down its old read key the moment it asks for an update breaks its own connection. ## What an observer sees Because the message travels after the handshake, it is inside a protected record like everything else, so the outer `opaque_type` reads `application_data(23)`. On the wire a rotation is a short protected record indistinguishable in kind from a subtitle line. The event is invisible to anything on the path that is merely watching record headers. ## Operating it on a continuous feed - **Rotate on a budget, not a hunch.** Count records or bytes written under the current key and rotate at a threshold well below any limit, rather than on wall-clock time alone. - **Rotate before the counter forces you to.** Reaching the counter's maximum with no rotation leaves only one legal option: closing the connection, which on a live feed is the outcome rotation existed to avoid. - **Expect asymmetry.** An ingest connection may write millions of records in one direction and a handful in the other. The busy direction will need rotating long before the quiet one, and that is exactly what a one-directional mechanism is for. - **Do not confuse this with resuming or re-authenticating.** A key update changes record-protection keys only; it re-runs no negotiation, proves no identity, and changes nothing that was agreed during the handshake. ## The failure to recognise If one side installs its new key too early - before its `KeyUpdate` has been sent, or before the peer's has arrived - the very next record it handles is opened or sealed under keys the other side is not using. The symptom is an immediate fatal integrity failure on a connection that was healthy a millisecond earlier, at a point unrelated to anything the application did. On a long-lived feed that presents as periodic unexplained disconnects, which is a much harder thing to diagnose than a handshake that never completed.

  • Under which key is the KeyUpdate message itself protected?
    The one currently in force. The sender seals the message with its existing write key and switches to the new one only afterwards, so the receiver can always open it with the key it already has. Switching first would make the message unreadable to the peer.
  • After sending KeyUpdate with update_requested, why do records still arrive under the peer's old key?
    Because the request rotates nothing in the peer's direction by itself. Records already in flight were protected before the peer saw the request, and the peer's direction changes only when it sends its own `KeyUpdate`. A receiver must keep accepting the peer's current key until that message arrives.
  • What forces a rotation rather than merely permitting one?
    The 64-bit write counter must never wrap, so an endpoint approaching its maximum either rotates or closes. On top of that, the negotiated authenticated-encryption algorithm may limit how much may be protected under one key, and a busy long-lived connection can reach that limit first.

saying these in an interview costs you the question

  • Thinks rotating traffic keys requires a fresh handshake
  • Says one KeyUpdate changes both directions at once
  • Answers an update_requested with another update_requested
  • Claims the KeyUpdate message is protected under the new key
  • Believes the record counter keeps running after a rotation
  • Treats a key update as re-authenticating the peer