skip to content

Where do the trust boundaries sit in a retrieval- and tool-using LLM assistant?

level: seniorimportance: must knowfreq 62%

answer

  1. the HTTP edge is not the line
  2. trust follows the writer, not the store
  3. context window is one zone
  4. the boundary is at the privileged sink
  5. the model enforces nothing

basics

~20 s

Not at the network edge. Everything that lands in the context window is one untrusted zone, whatever its source. The boundaries that matter sit between that context and any privileged action, and between each writer and each data source the system ingests.

solid answer

~50 s

The instinct is to draw the boundary around the user and treat internal corpora and tool results as inside. That is the mistake. Trust follows **who can write**, not where the bytes are stored, so a vendor manual in your own index is third-party text and a tool result is third-party text the moment the tool touches the outside world. Model the context window as a single untrusted zone and put the boundary where enforcement can actually happen: at every privileged sink the request can reach — a write API, a shell, a database, an outbound HTTP call, a rendered link. On a plant-maintenance copilot reading PLC streams, supplier manuals and work orders, the crown-jewel boundary is the SCADA write path, and it must be enforced in deterministic code with its own identity and authorization, because the model itself is not an enforcement point.

go deeper

for a junior

Know that anything placed in the prompt — user text, retrieved documents, tool results — is untrusted, and that authenticating the user does not make the retrieved content safe.

for a middle

Explain trust as a property of who can write to a source, and place enforcement at the tools and sinks rather than in prompt wording. Be able to classify a corpus, a database and a tool result.

for a senior

Walk a real data flow end to end, mark each writer and each privileged sink, and state which paths reach a crown-jewel action after consuming untrusted content. Say plainly that the model is not an enforcement point.

for a principal

Own the cross-team consequence: ingestion paths have owners outside your org chart, and the crown-jewel sink usually belongs to another function whose sign-off you need before shipping.

## The boundary people draw, and why it is wrong Asked to mark trust boundaries on an LLM application, most teams draw one line at the HTTP edge: outside is the internet and the user, inside is the app, the retrieval index, the tool layer and the model. It mirrors how they diagram every other service, and for a stateless CRUD API it is roughly right, because in that system nothing the user sends is ever interpreted as a command by a downstream component. An LLM application breaks that assumption. The context window is fed from many sources, and everything in it is available to the model as instruction. Authenticating the caller at the edge tells you who initiated the request; it tells you nothing about who wrote the 4,000 tokens of retrieved manual that arrived three hops later. ## Trust follows the writer The usable rule is that a piece of data inherits the trust level of **whoever can write it**, not of the store it sits in or the network zone it travels through. Apply that to a plant-maintenance copilot that reads PLC sensor streams, supplier equipment manuals and the work-order database, and can schedule a line shutdown: - The vendor manual corpus is **untrusted**. Suppliers upload PDFs; nobody in your org authors them and nobody reads all of them. - The work-order database is **semi-trusted**. Writers are authenticated employees and contractors, but the fields are free text and the population is large, so an entry is attributable but not vetted. - The PLC telemetry stream is **machine-authored but attacker-influenceable** — values and labels can be steered by anyone who can affect the plant floor or the historian. - The SCADA write path is the **crown jewel**: the one place where the system's output stops being text and becomes physical consequence. Once that is written down, the map is no longer "user versus system". It is a set of sources of differing provenance flowing into a shared context, and one high-consequence sink. ## Where the boundaries actually are **Between each writer and each ingested source.** Every path that puts bytes into a corpus, memory store or tool result is a boundary crossing with a named owner on the far side. This is the boundary teams most often omit, because the ingestion job looks like plumbing rather than an attack surface. **Between the context window and every privileged sink.** The context is one zone; the sinks are the boundary. Concretely: any tool that writes, spends, sends, deletes or schedules; any string handed to a shell, a SQL engine, a template renderer or a browser; any outbound request that could carry data off the premises. Enforcement at that line must be deterministic code with its own view of the caller's identity and scope, because it is the only part of the system whose behaviour does not depend on how persuasive the surrounding text was. **Between the model's output and any interpreter.** Model output is untrusted input to whatever consumes it. If a downstream component parses it as markup, code or a command, the boundary is that parser, and the control is validating and escaping there — the same discipline you would apply to a form field. **Between tenants, sessions and agents.** Anywhere context or memory can cross from one principal to another, there is a boundary. A memory record written under one user's session and recalled under another's is a crossing even though nothing left the process. ## The model is not a boundary The single most consequential line in this exercise: **the model is not an enforcement point.** It cannot reliably separate instruction from data, and defenses that ask it to — system-prompt admonitions, delimiters, spotlighting, instruction-hierarchy training — are probabilistic. Adaptive attackers have broken published defenses of this kind at very high success rates. Frontier models may be under a percent on a single attempt and several percent across a hundred adaptive attempts, which is fine as risk reduction and useless as a boundary. Whatever must not happen has to be impossible in code that runs outside the model. A corollary that surprises people: giving the model a *stronger* system prompt does not move any boundary. Adding a deterministic check in front of the shutdown-scheduling call does. ## Turning the map into findings For each boundary, record three things: what crosses it, who owns the far side, and what enforcement runs at the crossing. That is what makes the map actionable. "Retrieved manual text crosses into the context window; owner is the supplier-onboarding team; enforcement is none today" is a finding with an owner. "We use RAG" is not. Then check reachability: for every privileged sink, list the request paths that can trigger it and mark which of them have consumed untrusted content. A sink reachable only from paths that touched nothing external is a genuinely different risk from one reachable after reading a supplier PDF. That reachability question — not the model's benchmark refusal rate — is what determines whether the design is defensible. ## Common failure signals Teams that have not done this exercise say things like "the corpus is internal so it is trusted", "the tool is ours so its output is safe", "we authenticate users so the input is fine", and "the system prompt forbids that". Each of those sentences names a boundary that does not exist.

  • How would you decide the trust level of a tool's output?
    By asking who ultimately authored the bytes it returns. A tool that reads a fixed internal table returns data your org controls; a tool that fetches a URL, reads a ticket body, queries a shared inbox or shells out returns third-party content wearing your tool's name. Trust level is inherited from the furthest upstream writer, not from the fact that your team implemented the function.
  • Does putting retrieved documents in a separate message role create a boundary?
    No. Role separation, XML tags and delimiters are formatting conventions inside one context window; the model can be argued across them, and adaptive attacks do exactly that. They are worth having as risk reduction and as provenance labelling for your own logging, but never record them as the control that closes a finding on a high-consequence path.
  • Where does a memory store sit on this map?
    It is both a sink and a source, which is why it deserves its own boundary. Content written during one session becomes ingested context in later ones, so a payload planted once persists and can cross session or user scope. Model the write path (who can cause a write, and with what content) separately from the read path (which principals can recall it).

Treat the context window like a mailroom that accepts letters from anyone: you do not decide whether to act on a letter by checking which desk it is sitting on, you decide at the door of the room that holds the keys.

saying these in an interview costs you the question

  • Draws the only trust boundary at the HTTP edge
  • Calls an internal corpus trusted because of its network location
  • Treats tool output as trusted because the team implemented the tool
  • Names the system prompt as the enforcement point
  • Ignores model output as untrusted input to downstream parsers

context