A planted reference reaches an assistant's answer and nobody clicks it — how does data leave?
answer
- who actually issues the request
- the model opens no sockets
- displaying is an action, not a pause
- the client resolves references while drawing
- no decision point means no click to gate
basics
~20 sDisplaying is enough. If the answer carries a reference the viewing client resolves on its own — a preview it expands, a resource it embeds — the request goes out while the answer is being drawn. Nobody clicks.
solid answer
~50 sTwo things get merged here: the model producing text, and a client displaying it. The model writes no packets — every byte that leaves is written by something downstream. Many display surfaces resolve references while they draw: a bare link expanded into a preview card, an embedded resource pulled by a document viewer, a thumbnail generated for a digest. That request goes out while the answer is being shown, so the click a reviewer is counting on was never in the loop. An output screen applied after generation sees a bland sentence with a reference in it and scores it harmless — which is accurate, because nothing in the text is harmful; the harm is what a downstream program does with it. The channel is a property of the rendering surface, not of the model: the same answer is inert in a console and a network request in a rich client.
go deeper
Be ready to say plainly that the model writes no packets and the display surface does. If the client resolves a reference while drawing the answer, the request happens with no click at all.
Explain which component issues the request and when: rendering resolves references as part of layout, before any human decision. Note that a post-generation output screen scores prose and cannot see a downstream renderer's behaviour.
Show that the channel is per-surface and per-version, so the question 'is this exploitable' is answered by testing each display client with a fixture, not by reasoning about the model. Say what an arriving request does and does not prove.
Own the framing that egress capability lives in surfaces your team may not ship or inventory, and that no assurance about them holds without a named owner and a re-test cadence.
## Two events, not one Picture a data-analysis assistant that answers questions over database rows and spreadsheets. One column is free text — a notes or description field an outsider filled in through a public form. The assistant retrieves the row, quotes that field back inside its answer, and the answer is displayed somewhere. People treat *the model answered* and *the answer was displayed* as a single event. They are two events, owned by different software, and only the second one touches the network. The model emits tokens; it opens no sockets. Every byte that leaves the organisation is written by a program downstream of the model. ## What a rendering surface does while it draws Display surfaces are not passive. Many of them resolve references as part of drawing the content: a chat surface expands a bare link into a preview card by fetching it; a mail or document viewer pulls embedded resources so the layout is complete; a digest generator builds thumbnails; a dashboard tile loads a remote asset a cell refers to. None of that requires a human gesture. The request is issued because the surface decided the answer contains something it should resolve in order to show the answer properly. That is the entire mechanism. If the answer carries a reference of a form the viewing client resolves, and the reference points at a host the attacker controls, then *displaying* the answer performs the fetch. ## What the construction gets past Two things, and they fail for different reasons. **The human.** The reviewer's mental model is that a link is dangerous only if someone clicks it. A click is a control that assumes a decision point. Here there is no decision point: the request happens before anyone has decided anything, often before anyone has read the sentence. A control that was never in the loop cannot be credited with stopping anything. **The output screen.** A screen applied after generation scores the answer text — is it harmful, does it disclose, does it read like an attack. What it sees is a bland summary with a reference in it, and it scores that harmless. The score is correct. Nothing in the text is harmful. The harm is in what a different program does with the text a moment later, and that behaviour is not a property of the prose. ## What actually leaves Two different things, and interviewers want them separated. First, a **signal**: the request itself. Its arrival establishes that a particular surface dereferenced a model-authored reference at render time. That single bit is often the whole objective of a first attempt, because the attacker cannot see which of the unknown clients is live. Second, whatever the model wrote **into the reference's own structure** — path segments, parameters — data smuggled inside something the client treats as an address rather than as content. Keep the direction of every claim straight. A received request proves the surface resolved the reference. It does not prove a person read the answer, it does not prove the model reliably produces the reference, and it does not prove that other surfaces behave the same way. ## What it costs the attacker More than it looks. Whoever writes the field is writing for a renderer they cannot see, cannot version and cannot test. Rendering behaviour differs per surface and per client release; sanitisers differ; some surfaces show answers as plain text and resolve nothing at all. A reference form that is live in one client is inert in the next. The feedback loop is also entirely one-way. Nothing comes back except a request that may never arrive, and its absence is unreadable: the field may never have been quoted, the answer may never have been displayed, the client may have rewritten the reference. That blindness is the defining cost of this class, and it is why the attacker prefers surfaces likely to be *numerous* over surfaces likely to be *interesting*. ## Where it stops working - The surface shows answers as plain text and resolves nothing while drawing. - The reference is a form the client displays rather than fetches. - The answer never quotes the field — the pipeline paraphrases, or the row is never retrieved. - A resolved preview is cached: the fetch happens once for that reference and never again, which also means a second observation attempt sees nothing. ## The sentence to have ready The egress channel is a property of the rendering surface, not of the model. The same answer text is inert in one place and a network request in another — and nobody clicked anything either way.
- The output screen scored the answer as harmless. Was it wrong?No. It scored the text, and the text is a bland sentence with a reference in it. Nothing about it is harmful as prose. The event that matters happens after scoring, in a different program, and is a property of that program's rendering behaviour rather than of the words. Reading the screen's verdict as evidence that nothing left is the mistake — it proves the text scored below a threshold, not that no request was issued.
- The team says answers are plain text, so there is nothing to fetch. What would you check?Which surfaces actually display the answers. Plain text on the server says nothing about the client: a chat surface can expand a bare address into a preview, a mail digest can build a thumbnail, a mobile client can auto-link. The property to establish is per-surface and per-version — what does this client resolve while drawing — and it can be checked with a fixture answer, without the model in the loop at all.
- What does a single arriving request tell the person who planted the field?That one surface, once, resolved a model-authored reference while displaying an answer, and therefore that the field text reached that answer. It does not tell them who read it, how often the model quotes the field, or whether any other client behaves the same way. Silence tells them almost nothing, because the field never being quoted and the client never resolving look identical from the outside.
A read receipt in an email: you never opened an attachment or pressed a button, but the client fetched a remote element to lay the message out, and the sender learned you were there.
saying these in an interview costs you the question
- Says nothing leaves unless the user clicks the link
- Thinks the model itself opens the network connection
- Treats a passing output screen as proof no request was issued
- Assumes plain-text answers cannot trigger a client-side fetch
- Believes one surface's behaviour describes all the others