How should an LLM app propagate cancellation when a user closes the tab mid-stream?
answer
- a chain, not an event
- browser, server, provider — three hops
- the handler does not stop by itself
- the last hop is the one that bills
- nothing errors when it is broken
basics
~20 sCancellation must travel all three hops: the browser aborts the request, the server detects the closed connection, and the server aborts its own upstream call to the model. Break any hop and generation continues, billed, with nobody watching.
solid answer
~50 sTreat it as a chain, not an event. The browser aborts the in-flight request — typically an `AbortController` signal passed to `fetch`, fired on page unload or navigation. The server must notice that its client is gone, which most frameworks expose as a request-cancelled or connection-closed signal but do **not** act on by default: the handler keeps running unless you wire the signal through. Finally the handler must abort the upstream model call, because the provider goes on generating until the request is closed. Miss that last hop and you pay for output no one will read and hold a serving slot while doing it. Around the mechanics, decide the semantics: whether to persist the partial output as a draft before aborting, what goes in the session log, and how a retry avoids appending a second overlapping answer to the first.
code
javascript · 11 linesfunction streamAnswer(payload) {
const controller = new AbortController();
const stop = () => controller.abort();
window.addEventListener("pagehide", stop);
return fetch("/api/chat", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify(payload),
signal: controller.signal,
}).finally(() => window.removeEventListener("pagehide", stop));
}go deeper
Know that a user closing the tab does not automatically stop the model — the request has to be actively cancelled, or tokens keep being generated and billed.
Trace the chain: browser aborts the request, the server detects the closed connection, and the handler aborts its upstream call. Explain that most frameworks surface the signal but do not act on it.
Show that you would verify propagation end to end and decide the semantics — whether partial output is persisted, what enters conversation history, and how retry replaces rather than appends.
Own it as a cost and capacity control: abandoned generations consume paid tokens and serving slots invisibly, so cancellation correctness belongs in your reliability tests and in how spend is attributed.
## Three hops, any of which can silently fail When a user closes the tab mid-generation, the useful mental model is a chain of cancellations, each of which must be explicitly wired: 1. **Browser to your server.** The page is going away, so the request must be aborted. In practice you pass an `AbortController` signal to `fetch` and call `abort()` on unload or navigation. If you do nothing, behaviour varies by browser and by how the request was made — the connection usually dies with the page, but you should not leave it to chance, and some request modes deliberately outlive the page. 2. **Your server noticing.** The server sees a closed connection or a cancelled request context. Almost every runtime exposes this — a cancellation signal on the request, a close event on the connection — and almost none of them stop your handler for you. A handler that ignores it keeps looping over the upstream stream, writing into a socket nobody reads. 3. **Your server to the provider.** This is the hop that costs money. The provider is generating tokens until its request is closed. Aborting your own outbound call is what actually stops generation and releases the serving slot. The failure mode people miss is hop three surviving hop two: the server correctly detects the disconnect, logs it, returns from the handler — and the upstream call was made with a signal nobody cancelled, so generation runs to completion in the background. ## Why this matters beyond tidiness A long answer that nobody reads is still generated output: you are billed for the tokens produced before the abort, and every token after the user left is pure waste. Under concurrency the effect compounds — abandoned generations occupy capacity that live users are queueing for. On a page where users routinely reject the first sentence and retry, unpropagated cancellation can be a meaningful fraction of total spend. ## Decide the semantics, not just the mechanics Cancellation raises questions the plumbing does not answer: - **Do you keep the partial output?** For a chat turn, usually discard it. For a drafting tool, users are often glad to keep 300 words, so persist what arrived before aborting and mark it incomplete. Decide deliberately, because the natural implementation — abort and lose everything — is silently a product decision. - **What goes in the session log?** If the conversation history records a truncated assistant turn, the next turn is conditioned on a half-finished message. Either store it explicitly flagged as truncated, or do not store it at all; the wrong answer here produces strange follow-up behaviour that is hard to trace back. - **How does retry behave?** A retry after a partial render must replace, not append. Concatenating two independent generations produces duplicated or contradictory text. ## The client side of an aborted stream The same discipline applies when a stream dies mid-sentence for any other reason — a network drop, an upstream error at token 300. Text is already on screen and you cannot unsay it. The defensible options are to mark the fragment as incomplete with a clear affordance to continue or retry, or to roll the fragment back to the last committed boundary. What is not defensible is leaving a truncated sentence looking like a finished answer, because the user cannot tell the difference between "the model stopped here" and "the model meant this". If your renderer commits at sentence or block boundaries rather than per token, rollback becomes a clean operation rather than a guess. ## Testing it This is straightforward to verify and almost never verified. Start a long generation, kill the client, and check three things: that the upstream call ended promptly rather than at its natural completion, that your token accounting shows the truncated amount, and that no orphaned work continued behind the handler. Because nothing errors when propagation is broken — the user is gone, no alert fires, the dashboards look normal — the only way you learn about it is by looking, or eventually by the bill. ## Adjacent traps Background work kicked off by the request needs its own cancellation; aborting the model call does not stop a tool invocation or a follow-up job the handler already launched. And a keep-alive or fire-and-forget request mode that is designed to outlive the page will happily defeat hop one, so check the options you are passing rather than assuming page teardown cleans up.
- Your server logs the disconnect correctly but the provider bill says generation ran to completion. What is broken?The third hop. The server noticed its client was gone and returned, but the outbound call to the model was issued without a cancellation signal wired to that event, so it kept streaming into nothing. The fix is to derive the upstream request's cancellation from the inbound request's signal, so closing one closes the other.
- Should the truncated assistant message be written to conversation history?Only deliberately, and only flagged. If you store a half-finished turn as if it were complete, every subsequent turn is conditioned on a message that stops mid-sentence, which produces odd continuations that are hard to diagnose. Either mark it explicitly as truncated so the next prompt can account for it, or drop it and treat the turn as never having happened.
- How should the UI handle a stream that fails at token 300, mid-sentence?Never leave the fragment looking like a finished answer. Mark it visibly incomplete and offer continue or retry, or roll back to the last committed boundary — a sentence or block — and show the failure. Committing rendered output at boundaries rather than per token makes that rollback exact instead of a guess.
saying these in an interview costs you the question
- Assumes closing the tab stops generation at the provider
- Logs the client disconnect but never aborts the upstream call
- Leaves a truncated sentence on screen as if it were the answer
- Retries after a partial render by appending to the fragment
- Stores a truncated assistant turn in history as a complete message