When should an MCP tool mint a state handle instead of taking the state in every call?
answer
- both are legal, so it is a cost question
- context tokens versus server storage
- what should not be in a transcript
- live objects cannot be serialised
- the handle brings obligations with it
basics
~20 sMint a handle when the state is large, sensitive, or meaningless to the model — a handle keeps it out of the context window. Re-send it in arguments when it is small and self-describing, which avoids owning storage, expiry and cleanup for every workflow.
solid answer
~50 sBoth options keep requests self-contained, which is what MCP 2026-07-28 requires, so the choice is an engineering tradeoff rather than a compliance one. **Re-sending in arguments** is the cheaper default: the server owns nothing between calls, there is no expiry to tune, and a re-issued request after a dropped stream carries everything it needs. It fails when the state is large (it burns context tokens on every call), sensitive (it sits in the transcript), or opaque to the model (a serialised cursor the model can corrupt). **Minting a handle** fixes those, at the price of real ownership: shared storage every node can reach, a TTL and cleanup policy, a clear expired-handle error, and per-use authorization since possession of a handle MUST NOT be treated as authentication. My rule of thumb is to re-send by default and mint only when the state is genuinely server-side — an open cursor, a scratch workspace, a partially built upload — and then keep the handle opaque, short-lived and narrowly scoped.
go deeper
Know that both shapes exist and that a tool argument is the normal way state travels; a handle is what you use when the real state has to stay on the server.
Be able to name the immediate costs on both sides: repeated arguments consume context on every call, while handles require the server to store, expire and clean up what they point at.
Argue from the workflow — size, sensitivity, whether the state is a live object, and retry behaviour after a dropped stream — and attach the security obligations to the handle option rather than treating it as free.
Set the convention rather than deciding per tool: one handle format, one lifetime policy, one backing store, one expiry error shape, ownership enforcement everywhere, and a clear rule for when a tool is allowed to mint at all.
## Framing the decision Since revision 2026-07-28 made MCP stateless, both designs are legal and both satisfy the requirement that every request be self-contained: arguments are self-contained by definition, and a handle is itself just an argument. So the question is not compliance, it is who bears the cost — the model's context window and the wire, or the server's storage and lifecycle machinery. ## The case for re-sending state in arguments - **The server owns nothing.** No store, no TTL, no cleanup job, no orphan reaper, no cross-region replication question. This is the largest hidden cost of handles and it is easy to underestimate. - **Retries are trivially correct.** When a response stream drops, the client re-issues the call with a new JSON-RPC id; if all context is in the arguments, the retry is complete by construction and cannot reference state that has since expired. - **The workflow is legible.** A reviewer reading the transcript sees exactly what the call was about, which matters when a human is approving a destructive operation. - **It scales to zero.** Serverless and edge deployments need no backing store at all for these tools. It breaks down when the payload grows. Every argument is tokens in the model's context on every call, and a large blob repeated across a ten-step workflow is both expensive and a source of transcription errors when a model regenerates it rather than copying it. ## The case for minting a handle - **Size.** A query result set, a document, a file being assembled — these belong on the server, referenced by a short id. - **Sensitivity.** Anything you do not want sitting in a transcript should not be an argument. A handle keeps the value server-side and puts only an opaque reference in context. - **Genuinely server-side objects.** An open database cursor, a transaction, a sandbox process, a partially uploaded file are not serialisable into arguments in any honest way. The handle names a live thing. - **Integrity.** State the model cannot see is state the model cannot corrupt or a prompt injection cannot rewrite. A serialised pagination cursor in arguments can be tampered with; an opaque id checked against server records cannot be repurposed as easily. The costs are equally concrete. You now own storage reachable from every node — keeping it in process memory recreates session affinity that the load balancer does not know about, and it will fail on scale-in and rolling deploys. You own a TTL and the error behaviour when it lapses. You own per-use authorization, because 2026-07-28's State Handle Hijacking rule says possession MUST NOT be treated as authentication, so every use must be checked against the credential on the current request. And you own the cleanup path, including the common case where the client simply walks away mid-workflow — with a closed stream read as cancellation, abandonment is normal, not exceptional. ## A decision rule Re-send in arguments by default. Mint a handle when at least one of these is true: 1. The state exceeds a size where repeating it in context is wasteful. 2. The state is sensitive and must not appear in a transcript. 3. The state is a live server-side object that cannot be serialised. 4. Model tampering with the state would be harmful. And when you do mint, hold yourself to the whole contract: opaque and unguessable value, short TTL with an actionable expiry error, narrowest possible scope, ownership bound to the validated principal and re-checked on use, shared storage, and a repeat-safe design because a dropped stream will bring the same handle back under a new request id. ## The organisational angle This is not a per-tool decision to make thirty times independently. Set the convention once for a server or a platform: what a handle looks like, how long it lives, which store backs it, how expiry is reported, and how ownership is enforced. Teams that decide ad hoc end up with four handle formats, three expiry behaviours and one tool that quietly treats the handle as a credential. A hybrid also works well — a handle for the bulk object, plain arguments for the small parameters that vary per call — and is usually the shape that reads best in a transcript. ## What interviewers are listening for That you know both are protocol-legal, that you can name the specific costs on each side rather than reciting "stateless is better", and that you attach the security and lifecycle obligations to the handle option rather than treating it as free.
- Give an example where the hybrid is clearly right.A data-analysis server: `open_dataset` returns a `datasetId`, and each `run_query` call takes that id plus the query text and row limit. The bulk state — the loaded dataset and its schema — stays server-side, while the parameters that actually vary per call stay visible in the transcript, which keeps the calls readable for a human reviewing what the agent did.
- How does the choice change if the server runs serverless with no persistent store?It pushes you hard toward arguments, since a handle needs storage reachable from every invocation and instance memory does not survive. If you still need a reference, a short-lived signed value that carries the small state itself removes the store dependency — but accept the consequences: it cannot be revoked before expiry without a denylist, and it must contain nothing sensitive, because it travels through model context.
- Does putting state in arguments risk anything security-wise?Yes — everything in arguments is in the transcript and in whatever logs capture tool calls, and it is visible to the model, so a prompt injection can attempt to alter it between steps. That is precisely why sensitive or integrity-critical state belongs behind a handle. The tradeoff runs both ways; neither option is the safe one in general.
- How do you size the TTL on a handle?To the workflow, not to a round number. Measure how long real multi-step sequences take, take a generous upper percentile, and add margin for a human approving a step in the middle. Then make expiry a clear, distinguishable tool error so the model re-mints instead of retrying blindly, and monitor the expired-handle rate as the signal that your window is too tight.
saying these in an interview costs you the question
- Treats handles as free because the protocol allows them
- Says arguments are wrong because MCP is stateless
- Ignores that arguments consume model context on every call
- Puts sensitive state in arguments to avoid owning storage
- Mints handles without a TTL or a cleanup path