How do you decide between hot-reloading TLS key material through GetCertificate and rolling restarts fleet-wide?
answer
- what a restart costs this service
- connection lifetime changes the answer
- code that runs quarterly is broken code
- the overlap window is the real invariant
- alert on remaining validity, not reload success
basics
~20 sDecide on what a restart costs you: long-lived connections and per-process resumption state, versus the maintenance cost of swap code that runs rarely. Pick one posture per fleet, exercise it continuously, and alert on remaining validity rather than on reload success.
solid answer
~50 sThe two postures are: reload in place, where a `GetCertificate` callback returns whatever key pair is currently published and a reload step swaps it atomically; and rolling restart, where renewal writes new files and the existing deploy path replaces processes. Restarting reuses machinery you already trust, but it drops every established connection, discards per-process session resumption state and re-warms pools — which is cheap for short request-response traffic and expensive for long-lived streams. Hot reload is invisible to live connections but adds a concurrency-sensitive code path that runs once a quarter, and code that runs once a quarter is code that is broken. My call is usually: one reviewed internal package owning the callback, the atomic swap and its race test; every service uses it; renewal runs often enough in staging that the path is genuinely exercised. And whichever posture you choose, the retiring material stays acceptable through an overlap window, and the alert that matters fires on remaining validity, not on the last reload's exit code.
go deeper
Understand that new key material on disk does nothing by itself: either the process reloads it or the process is replaced, and both are deliberate choices someone made.
Be able to compare the two mechanisms concretely — what a restart drops, what the reload callback preserves — rather than asserting that one is simply more modern.
Argue from measurements of your own service: connection lifetime, warm-up cost, renewal cadence, and the overlap window the slowest consumer needs. Say what you alert on and what a failed reload must not do.
Own the posture across the fleet: one implementation rather than forty, a cadence that keeps the path genuinely exercised, and a clear line on when TLS termination should move out of the services entirely.
## What is actually being decided Every service that terminates TLS itself has to answer one question on a schedule: when new key material exists, how does the running process start using it? There are two defensible answers and they are not equally good in every service, which is why this is a posture the owner sets rather than a rule that falls out of the language. **Reload in place.** The server's `tls.Config` uses a `GetCertificate` callback that hands back the currently published `tls.Certificate`. A reload step — a timer, a signal handler, a filesystem watch — parses the renewed files and publishes them with an atomic swap. Connections already established carry on unaffected; the next handshake gets the new chain. **Rolling restart.** Renewal writes the files and then something replaces the processes, one at a time, using whatever machinery already ships new binaries. A third answer exists — terminate TLS somewhere shared and let the services speak plaintext or a separate internal identity behind it — and it is often the right one at fleet scale, because it collapses N rotation stories into one. Noting that option is part of answering well; it is also the answer that moves the decision out of the service team's hands, which is exactly why it needs to be made deliberately. ## What a restart actually costs The honest input to this decision is a measurement, not a preference. - **Connection lifetime.** If your traffic is short request-response exchanges through a balancer that drains, a restart is nearly free — you already do it on every deploy. If clients hold long-lived streams for minutes or hours, a restart severs all of them at once and every client reconnects in the same second. The reconnect storm, not the restart, is what hurts. - **Per-process state.** A fresh process starts with an empty connection pool downstream, a cold cache, and — unless you have configured them explicitly — freshly generated session ticket keys, so resumption for that replica's clients stops working until they re-establish. - **Frequency.** Certificate lifetimes have been getting shorter, and a renewal cadence measured in days makes "just restart" a materially different proposition from a yearly one. ## What hot reload actually costs The swap is not hard, but it is a concurrency-sensitive path: the callback runs on every handshake goroutine while the reload writes. Getting it right means publishing complete, immutable material through an atomic and never mutating it afterwards, and proving it with a test that renews under load with the race detector on. That is a modest amount of code — and it is code whose failure mode is that it does nothing at all until the day it must work. That is the real argument, and it is an organisational one. A rotation path exercised once every ninety days is untested. If you build it, you must **exercise it continuously**: rotate on a short cycle in staging, and rotate in production far more often than expiry demands, so a broken reload surfaces on a Tuesday afternoon rather than at the moment the old certificate dies. If you are not willing to do that, the rolling restart — a path you exercise on every deploy — is the more reliable choice even though it is the cruder one. ## The invariant both postures share Whichever you pick, the same rule governs the changeover: **the retiring material must stay acceptable until everything that depends on it has moved.** Concretely that means the overlap window is sized by the slowest consumer, not by the swap. Live connections keep the chain they negotiated. Peers that pinned or cached the old material need time to refresh. Anything that verifies rather than presents — a client trusting a root, a service accepting tokens signed by a retiring key — must accept both old and new for the whole window, and the ordering is always the same: **distribute the new material for verification everywhere first, only then start presenting or signing with it, and retire the old one last.** Reversing those steps is how a rotation becomes an outage, and it is the failure that gets written up afterwards. ## What you own operationally - **Alert on remaining validity of the material the process has actually loaded**, not on the file on disk and not on whether the last reload logged success. A process happily serving a chain it loaded three months ago will pass a "reload succeeded" check right up to expiry. The threshold should sit far enough out that a human can intervene after the automation has had several chances to retry. - **A failed reload must be non-destructive.** If the new files do not parse, keep serving the previous pair and raise an alarm. Publishing a truncated file turns a renewal glitch into an immediate outage. - **Decide who owns the code.** Forty services each writing their own swap gives you forty chances at a data race and forty different failure behaviours. One internal package with the callback, the atomic swap, the reload trigger and the race test is the fleet-level version of this decision, and it is the part a platform team can legitimately overrule a service team on. - **Know what you would do at 3am.** If rotation automation fails and expiry is hours away, is the runbook "trigger a reload" or "deploy"? Both are fine; not knowing which is not. ## How I would summarise the call Short-lived connections, frequent deploys, few replicas: restart, and spend the saved complexity budget elsewhere. Long-lived connections, expensive warm-up, or renewal cadence measured in days: build the reload path, but only alongside a commitment to exercise it constantly and to own it in one place rather than forty.
- What should a certificate-expiry alert for a Go service actually be based on?The remaining validity of the material the process has loaded, reported by the process itself. A check on the file, or on whether the last reload logged success, passes happily while a server serves a chain it loaded months ago. Set the threshold far enough out that automation gets several retries and a human still has room to act.
- When would you not build the hot-reload path at all?When connections are short-lived, deploys are frequent and cheap, and the fleet is small. Then a rolling restart uses machinery you exercise daily, while the swap would be concurrency-sensitive code that runs four times a year and is therefore untested at the moment it matters. Complexity you cannot exercise is a liability, not resilience.
- How do you stop forty services each writing their own certificate swap?Put the callback, the atomic publish, the reload trigger and the renew-under-load race test in one internal package and make using it the default in the service template. It turns forty chances at a data race into one reviewed implementation, and gives the platform team a single place to change the posture later.
- In any rotation, which side moves first — the side that presents the new material or the side that accepts it?The accepting side. Distribute the new material for verification everywhere first, then start presenting or signing with it, and retire the old material last. Reversing that order means something is offered material a peer cannot yet validate, which is the ordering mistake that turns a routine rotation into an incident.
saying these in an interview costs you the question
- Treats rotation as a one-off script rather than a routine
- Assumes a restart is free because deploys are automated
- Ignores connections already established at swap time
- Leaves no overlap window for the retiring material
- Alerts on reload success instead of remaining validity
- Publishes material that failed to parse, breaking a working server
- Lets every service write its own swap implementation