When should a service allow handler hijacking, given it opts those connections out of platform guarantees?
answer
- who else depends on this connection behaving
- deploys and dashboards, not throughput
- you replace every guarantee you dropped
- isolate what the server cannot drain
basics
~20 sOnly when the workload genuinely needs a raw bidirectional stream and the team will fund the replacements: an explicit session bound, its own registry and close path, its own metrics, and isolation from routes whose deploy and drain behaviour must stay predictable.
solid answer
~50 sTreat it as a contract decision, not a capability question. A hijacked connection leaves `http.Server`'s world: its timeouts stop being applied, it is invisible to connection accounting because `StateHijacked` is terminal, and `Server.Shutdown` explicitly neither closes nor waits for such connections — so the drain behaviour the platform team promises for every deploy simply does not hold on that route. I first ask whether an ordinary streamed response the server still frames would do; that keeps every guarantee. If the requirement really is a raw bidirectional protocol, I approve it with compensations written down: a bounded session lifetime, a registry of live sessions with a close path wired into shutdown, session-level metrics owned by the upgrade code, and ideally a separate listener or deployment so a stuck session cannot slow the main service's rollouts. And I accept that the platform or on-call team can veto it, because the guarantees being dropped are theirs — if nobody will staff the compensations, the answer is no.
go deeper
Know that taking over a connection is not just a coding choice: it removes that connection from the server's control, so somebody has to decide whether the service is allowed to do it.
Be able to state which shared behaviours stop working — bounded requests, drain on shutdown, connection accounting — and why that makes the choice bigger than one handler.
Argue the alternatives first and then the compensations: what you bound, what you track, what you close, and how you prove during a deploy that the route drains.
Own the negotiation. Decide the policy, write the compensations into a decision record, isolate the route so ordinary deploys stay predictable, and accept that the team whose guarantees you are dropping holds a legitimate veto.
## Why this is a decision and not a technique Every other question about hijacking is mechanical: which values come back, what you must write, what stops being enforced. This one is different because the cost lands on people who did not write the handler. A hijacked connection stops behaving like the rest of the fleet, and the guarantees it drops are the ones an organisation builds its operations on. ## What is actually being given up - **Deploy drain.** `Server.Shutdown` is documented not to close or wait for hijacked connections; the caller must notify such connections and wait for them separately. A rolling restart that assumes "Shutdown returns when in-flight work is done" is now wrong for this route unless somebody wrote that extra code. - **Bounded work.** The server's read, write and idle timeouts no longer bound the session. Whatever bound exists is one your handler enforces, and "none" is the default. - **Accounting.** `StateHijacked` is a terminal `ConnState`, so the standard connection gauge never sees these sessions end, and per-request middleware records no meaningful status or latency. - **Capacity shape.** Long-lived sessions change the unit of load from requests per second to concurrent connections per instance, which changes autoscaling signals, load-balancer idle settings and the blast radius of an instance restart. ## The order I work through it **1. Is a raw stream genuinely required?** Many requirements phrased as "we need a persistent connection" are really "push updates as they happen", which an ordinary streamed response — still framed, still owned by the server — satisfies while keeping timeouts, drain and metrics. Polling with a short interval is unglamorous and keeps every guarantee. Only a genuinely bidirectional, non-HTTP protocol after the handshake justifies the handover. **2. Who pays for the compensations?** For each dropped guarantee there is a concrete replacement, and each is real work: a session lifetime and idle bound; a registry of live sessions with an owner for `Close`; shutdown wiring that notifies and drains those sessions before the process exits; session-scoped metrics and logs; and load tests measured in concurrent sessions rather than requests per second. If the team cannot commit to all of it, the honest answer is that the feature is not affordable yet. **3. Can the blast radius be contained?** The strongest structural move is isolation: put upgraded routes on their own listener, their own port, or their own deployment. Then the ordinary API keeps its fast, predictable rollout, and the long-session service gets its own — longer drain window, different scaling policy, different alerting. That converts an operational argument into a deployment topology decision, which is much easier to hold over time. **4. Who can say no?** The service owner proposes; the platform or on-call team can veto, and should be able to, because deploy drain and connection accounting are their guarantees to the whole fleet. Written compensations turn that from a standoff into a review. If the route ships without them, the first incident is someone on call at three in the morning discovering that a rolling restart hangs and the connection dashboard has been lying for months. ## What I would not accept as an argument "It works, we tested it" — the failures here are operational and appear during a deploy or an incident, not in a functional test. "We will add the metrics later" — the metrics are how anyone notices the sessions leaking, so later means never. "Only one endpoint does it" — one endpoint is enough to make a whole instance undrainable if a single session is stuck. ## The shape of a good decision record Name the protocol and why HTTP framing cannot carry it; state the session bound; name the owner of the close path and where shutdown reaches it; list the metrics and where they are graphed; state the deployment isolation; and name the condition under which the route would be withdrawn. That document is what lets a future on-call engineer, who has never read the handler, reason about a connection the server has no opinion about.
- What would you require in writing before approving a hijacking route?A bounded session lifetime and idle policy; a named owner for the connection's Close and the shutdown path that reaches it; session-level metrics and logs, since request middleware and the ConnState gauge will not see these sessions; a load test measured in concurrent sessions; and a deployment or listener boundary that keeps ordinary routes drainable.
- Which alternatives do you evaluate before accepting the handover?Anything that leaves the connection under the server's control: a streamed response the server still frames for server-to-client push, or short-interval polling when the update rate is low. Both keep the server's timeouts, drain behaviour and per-request metrics. Only a genuinely bidirectional non-HTTP protocol after the handshake justifies giving those up.
- How do you answer a platform team that wants to ban hijacking outright?Grant the premise — their drain and accounting guarantees really do stop at the handover — and negotiate on isolation and compensations rather than on the technique. A ban on hijacking in the shared request path, with an allowance for a separately deployed service that carries its own drain and metrics, usually satisfies both sides.
- What changes about capacity planning once a service carries upgraded sessions?The load unit becomes concurrent connections per instance rather than requests per second, so autoscaling signals, file-descriptor limits and load-balancer idle settings all have to be revisited, and an instance restart now interrupts sessions rather than requests, which lengthens the drain window you need.
saying these in an interview costs you the question
- Treats it as a purely technical capability with no owner outside the team
- Promises to add session metrics after launch
- Assumes Server.Shutdown will drain upgraded connections
- Puts long-lived upgraded routes beside latency-sensitive ones
- Reaches for a handover when a streamed response would do