Why do Go TLS servers behind one load balancer need tls.Config.SetSessionTicketKeys given identical keys?
answer
- why a resumed handshake stops resuming
- each process invents its own ticket key
- the client is not rejected, only slowed
- the key list is the overlap window
- first encrypts, all of them decrypt
basics
~20 sBy default each Go process generates and rotates its own session ticket key, so a ticket issued by one replica cannot be decrypted by another and the client pays a full handshake instead. SetSessionTicketKeys gives every replica the same key set.
solid answer
~50 sA TLS session ticket is the client's resumption state, encrypted under a key that only the server can read. Go generates that key per process and rotates it on its own schedule, which is fine for a single server and wrong behind a load balancer: a client resuming on a different replica presents a ticket nobody there can decrypt, so it silently falls back to a full handshake. `(*tls.Config).SetSessionTicketKeys(keys [][32]byte)` fixes that by handing every replica the same list. The first key encrypts newly issued tickets while every key in the list can decrypt, which is also how you rotate: prepend the new key, keep the retiring one in the list until outstanding tickets have expired, then drop it. It is safe to call on a running server, and it panics if the slice is empty. This shows up as latency and CPU, never as an error.
go deeper
Know that TLS can resume an earlier session using a ticket, and that resuming saves a round trip and expensive cryptography compared with a full handshake.
Explain that a ticket is server state encrypted under a server-held key, that Go picks that key per process by default, and what SetSessionTicketKeys changes about which keys encrypt and which decrypt.
Diagnose the symptom: no errors, just a collapsed resumption rate and handshake CPU tracking connection churn after a scale-out or a change in balancing, and lay out the prepend-then-retire rotation order.
Own the tradeoff — a shared ticket key buys resumption across replicas at the cost of a long-lived symmetric secret — and set the rotation cadence and distribution path rather than leaving one key pinned indefinitely.
## What a session ticket is A full TLS handshake costs a round trip and real asymmetric cryptography. Resumption avoids most of that: the server hands the client a **session ticket** — a blob containing the state needed to resume, encrypted and authenticated under a key the server holds. The client stores the opaque blob and presents it on a later connection; the server decrypts it and short-circuits the expensive part. This is stateless on the server's side: nothing is stored, everything the server needs comes back inside the ticket. That property is what makes the ticket key so important, because possession of the key is the *only* thing that lets a server read a ticket. ## The default in Go, and why it breaks across replicas When you do not configure anything, a Go server generates a session ticket key at random inside the process and rotates it periodically on its own. For one server this is a good default — it is one less secret to manage and the automatic rotation limits how much resumable traffic a single stolen key covers. Behind a load balancer it produces a subtle degradation. Replica A issues a ticket encrypted with A's key. The client comes back a minute later, the balancer sends it to replica B, and B cannot decrypt a thing. The client is not rejected: TLS falls back cleanly to a full handshake and the request succeeds. What you observe is a resumption rate near zero, extra round trips on connection setup, and asymmetric CPU that scales with connection churn rather than with request volume. Nothing in the logs says "resumption failed". The same thing happens to a single server across a restart, and across a rolling deploy, since a fresh process invents a fresh key. ## The API ```go func (c *tls.Config) SetSessionTicketKeys(keys [][32]byte) ``` Each key is 32 bytes. The list has an ordering rule that is the whole of the rotation story: - the **first** key is used to encrypt newly issued tickets; - **all** keys in the list may be used to decrypt a presented ticket. It is safe to call while the server is running, which is what lets you rotate without a restart, and it panics if `keys` is empty — pass at least one, and use `tls.Config.SessionTicketsDisabled` if what you actually want is no tickets at all. ## Rotating the ticket keys The ordering rule gives you exactly the overlap this leaf is about — the retiring key still verifying while the new one signs, expressed in ticket terms: 1. Generate a new 32-byte key and distribute it to every replica. 2. Once every replica has it, call `SetSessionTicketKeys` on each with the new key **first** and the retiring key still present. From this moment new tickets are encrypted under the new key and old tickets still decrypt. 3. Wait out the lifetime of tickets issued under the retiring key. 4. Call `SetSessionTicketKeys` again without it. The two ordering mistakes are the interesting ones. Promote the new key to first *before* every replica has the list and clients get tickets that some replicas cannot read. Drop the retiring key too early and every outstanding ticket is dead, giving you a burst of full handshakes — a CPU spike, not an outage, but a visible one on a busy fleet. How the key material reaches the replicas is a secret-distribution question rather than a Go one; what Go pins down is that the list must be identical and that the swap itself is a live call, not a restart. ## The tradeoff you are actually making Sharing one ticket key across a fleet trades a security property for a performance one. A session ticket key is a long-lived symmetric secret that can decrypt every ticket issued under it, and that undermines forward secrecy for resumed sessions: an attacker who records traffic and later obtains the ticket key can unwrap sessions that were resumed with it. Go's per-process auto-rotation is generous on this axis precisely because those keys are short-lived and never leave the process. When you pin a shared key, you take responsibility for rotating it on a schedule and for the blast radius if it leaks. "Set once at deploy and never think about it again" is the failure mode: the key ends up years old, in every config store, and shared with people who have long since left. ## What it is not The session ticket key is not the server's private key and has nothing to do with the certificate chain. Rotating one has no effect on the other; they have different lifetimes, different distribution stories and different consequences on compromise. Confusing the two is the most common mistake in this area — a candidate who says "we rotate the ticket keys when the certificate is renewed" has coupled two unrelated schedules for no reason. ## How you would notice Because the failure is invisible in error rates, you find it by measuring resumption: a counter of full versus resumed handshakes, or simply handshake CPU that tracks connection rate. A fleet that suddenly stops resuming after being scaled out, or after connections stopped being pinned to a replica, is the classic story.
- What does the order of the keys passed to SetSessionTicketKeys mean?The first key encrypts newly issued tickets; every key in the list is tried when decrypting one. So rotation is a two-step list edit: prepend the new key while keeping the retiring one so old tickets still resume, then remove the retiring key once tickets issued under it have expired. Ordering is the entire overlap mechanism.
- What is the cost of pinning one shared session ticket key across the fleet forever?It is a long-lived symmetric secret that decrypts every ticket issued under it, so it weakens forward secrecy for resumed sessions and its blast radius grows with age. Go's default per-process key is short-lived precisely to avoid that. If you share one, you owe it a rotation schedule and a distribution path you can revoke.
- How does a client behave when no replica can decrypt the ticket it presents?It performs a full handshake instead. There is no error, no alert and no failed request — just an extra round trip and the asymmetric cryptography resumption was meant to skip. That is why this shows up as connection-setup latency and CPU on a busy fleet rather than in error rates.
- Does rotating the session ticket keys have anything to do with renewing the server certificate?No. The ticket keys are symmetric secrets the server uses to encrypt resumption state; the certificate's private key proves identity during the handshake. They have separate lifetimes, separate distribution and separate consequences on compromise, and coupling their schedules only makes both harder to reason about.
saying these in an interview costs you the question
- Confuses session ticket keys with the certificate's private key
- Thinks a ticket that cannot be decrypted fails the connection
- Assumes replicas share ticket keys automatically
- Sets one shared ticket key and never rotates it
- Drops the retiring ticket key the moment the new one is added
- Calls SetSessionTicketKeys with an empty slice