skip to content

After a Go 1.24 upgrade, handshakes to one partner endpoint hang. How do you confirm the post-quantum group is the cause?

level: seniorimportance: should knowfreq 18%

answer

  1. the same binary talks to everyone else fine
  2. the hello outgrew one packet
  3. a hang, not an alert
  4. A/B the group in a two-dial harness
  5. scope the exception to one destination

basics

~20 s

Dial the same endpoint twice, once with Go's defaults and once with CurvePreferences pinned to classical curves. If only the classical dial completes, the larger X25519MLKEM768 ClientHello is the trigger. Fix it per destination, not fleet-wide.

solid answer

~50 s

Reproduce it in isolation first: a tiny harness that dials the failing endpoint twice from the same binary, once with a default `tls.Config` and once with `CurvePreferences` set to `[]tls.CurveID{tls.X25519, tls.CurveP256}`. If the classical dial succeeds and the default one times out, you have it — the `X25519MLKEM768` key share pushes the `ClientHello` past a single packet, and some servers and middleboxes mishandle one that arrives split across TCP segments. The tell is that it hangs rather than returning a clean alert, and only against that one peer. Then verify at fleet scale with a canary: compare handshake error counts, latency percentiles and counts of the negotiated group against the control. The fix is a scoped `tls.Config` for that destination with classical groups, an expiry date on the exception, and a ticket with the partner — not `GODEBUG=tlsmlkem=0` across the fleet.

go deeper

for a junior

Know that a bigger ClientHello is the visible consequence of the hybrid key exchange, and that comparing a working destination with a failing one is the first useful step.

for a middle

Be able to explain why the key share grew to roughly 1216 bytes, why that pushes the ClientHello across two TCP segments, and how pinning classical curves isolates the variable.

for a senior

Demonstrate the discipline: reproduce with an A/B dial, distinguish a peer that lacks the group from a peer that mishandles the larger hello, canary the change, and scope the exception to one destination.

for a principal

Own the rollout posture — canary before fleet, a published counter of negotiated groups as the adoption metric, and a standing rule that compatibility exceptions carry an owner and an expiry.

## Why this failure exists at all Go 1.24 made `X25519MLKEM768` the first key-exchange group offered on TLS 1.3 handshakes. The client's key share for that group is roughly 1216 bytes, against 32 for plain X25519. A `ClientHello` that used to be a few hundred bytes is now well over a kilobyte, which means it no longer fits in a single Ethernet frame and arrives at the server split across two TCP segments. That is entirely legal TLS. It is nevertheless the single most common source of breakage when this default lands, because a long tail of TLS terminators, load balancers and inspection middleboxes were written on the quiet assumption that a `ClientHello` arrives in one read. Some of them stall waiting for a message they think is complete; some drop the connection with no alert at all. The result on the Go side is a handshake that hangs until your dial timeout fires, not a crisp error naming a group. ## Confirming the cause Work from cheapest to most expensive. **1. Establish the pattern.** Is it one destination or many? A single partner endpoint failing while every other target works is already strong evidence, because the change is uniform on your side — your binary offers the same group to everyone. **2. A/B the group in a harness.** Write the smallest possible program that dials the failing host twice: once with a plain `tls.Config`, once with `CurvePreferences` pinned to classical groups. Two outcomes, one conclusion: - classical succeeds, default hangs → the hybrid group (almost certainly its size) is the trigger; - both hang → look elsewhere entirely; the timing of the upgrade was a coincidence. This takes minutes and produces the sentence you will put in the partner's ticket. **3. Confirm it is size, not group support.** A peer that merely does not *support* the group is not a failure case: TLS negotiates down, the handshake completes on a classical curve, and you get no post-quantum protection but no outage. So a hang is evidence of something mishandling the larger hello rather than something rejecting the group. This distinction matters because it changes who you escalate to and what you ask them to fix. **4. Check the benign skew before blaming anyone.** A Go 1.23 peer and a Go 1.24 peer share no post-quantum group — 1.23's pre-standard draft group was removed in 1.24 — so they fall back to classical X25519 and connect fine. "We are not getting post-quantum with this partner" and "we cannot connect to this partner" are different findings with different urgency, and conflating them wastes an incident. ## Confirming at fleet scale The harness proves the mechanism for one endpoint. To know what the upgrade did to everything else, use a canary rather than a rollout: put the new build on a slice of instances and compare it against the unchanged control on three signals. - **Handshake failures and timeouts per destination.** The signal that finds any other broken peer you have not noticed yet. - **Handshake latency percentiles.** The extra kilobyte and a half is usually invisible next to round-trip time, but a peer whose path is already at the edge of an MTU problem can show up as a p99 tail rather than an outright failure. - **Counts of handshakes by negotiated group.** Recent Go exposes the negotiated group on the connection state as `ConnectionState.CurveID`. Emitting a counter keyed by it turns "did the upgrade actually give us post-quantum" from a belief into a number — and it is the same counter that will later show a stale `CurvePreferences` somewhere in the fleet quietly suppressing the group. ## Choosing the fix There are two levers and they are not equivalent. `GODEBUG=tlsmlkem=0` disables the hybrid group for the whole process. It is the right tool for exactly one situation: it is 3am, many destinations are failing, and you want the blast radius closed now. It is the wrong tool for one broken partner, because it silently removes post-quantum protection from every other connection the service makes. A scoped `tls.Config` with classical `CurvePreferences`, used only for that destination, keeps everything else protected. That is the fix to ship. Attach two things to it: a comment naming the partner and the symptom, and a date to revisit — exceptions that outlive their cause are how a fleet ends up classical again without anyone deciding it should be. And open the ticket upstream: a stack that cannot handle a segmented `ClientHello` will break on the next protocol extension too, so it is worth their fixing rather than your permanently working around. ## The wrong moves Rolling the whole fleet back to Go 1.23 gives up every unrelated fix in the release to work around one peer. Raising the dial timeout converts a fast failure into a slow one and hides the problem. And blaming the group without the A/B leaves you arguing from a release note instead of from a reproduction.

  • Why does the failure look like a hang instead of a TLS alert?
    Because the peer or a middlebox is stalling on a `ClientHello` that arrived across two TCP segments rather than rejecting a group it dislikes. There is no protocol-level refusal to report, so your side simply waits until the dial timeout fires. A group the peer merely does not support would negotiate down cleanly instead.
  • A Go 1.23 client connects to your Go 1.24 server. What happens?
    The handshake succeeds on a classical curve. Go 1.23's pre-standard draft group was removed in 1.24, so the two share no post-quantum group and negotiate down. You lose the protection with that peer but nothing breaks — which is why a counter of negotiated groups is more informative than an error rate here.
  • Why not just set GODEBUG=tlsmlkem=0 across the fleet and move on?
    It is process-wide, so one misbehaving partner would cost you post-quantum key exchange on every other connection the service makes. Keep it as an incident lever when many destinations are failing at once; ship a `CurvePreferences` exception scoped to the one destination, with an expiry and an upstream ticket.
  • What would make you conclude the upgrade was not the cause?
    If the classical-only dial hangs too. That removes the key-exchange group from the picture entirely and points at routing, the peer's own outage, DNS, or a certificate or timeout change that shipped in the same deploy. The A/B harness is worth writing precisely because it can exonerate the upgrade.

saying these in an interview costs you the question

  • Blames the group from the release notes without reproducing
  • Disables post-quantum fleet-wide for one broken partner
  • Assumes a peer lacking the group causes handshake failures
  • Raises the dial timeout to make the symptom go away
  • Rolls the entire fleet back a Go release for one endpoint
  • Ships the exception with no expiry and no upstream ticket