In LLM security, how do direct and indirect prompt injection differ as risk categories?
answer
- two entry points, two owners
- who typed it versus who wrote it
- confused deputy through your own corpus
- every ingestion path widens one of them
- provenance, not prompt wording
basics
~20 sDirect injection arrives in the user's own turn, so the user is the adversary. Indirect injection hides in content the application feeds the model later — retrieved documents, tool results, memory — so a third party attacks through your data pipeline.
solid answer
~50 sThe model cannot tell them apart; the defender must. **Direct** injection means the hostile text is in the turn the end user typed, so the adversary is your authenticated user trying to reach data or actions they are not entitled to. The control is per-user authorization enforced outside the model, on the tool and data side. **Indirect** injection means the hostile text rode in on content your application fetched on the user's behalf — a supplier-uploaded manual, a work-order note, a web page, a prior memory record. The user may be innocent; the adversary is whoever can write to that source, and the risk is a confused deputy acting with your service's authority. That makes the register entry a list of ingestion paths with named owners, not one line. As of mid-2026 neither form is reliably stopped at the model layer.
go deeper
Be able to say plainly that injected instructions can arrive either in what the user types or inside documents and tool results the app fetches, and that the model cannot tell instructions from data.
Explain the mechanics: same context window, different entry point, different adversary. Name the confused-deputy shape of indirect injection and why the indirect surface grows with every ingestion path.
Show you can enumerate real ingestion paths for a live system, name a writer-owner for each, and state honestly that model-layer filtering is risk reduction rather than a boundary.
Own the framing that indirect injection is a supply-chain problem shared with teams and vendors outside your org, and drive the accountability question of who signs off on content entering the context window.
## What the split actually is Prompt injection is the class of failure where text reaching the model's context window is interpreted as an instruction rather than as data. Splitting it into *direct* and *indirect* is a statement about **where that text entered the system**. It is a threat-modeling distinction, not a technical one: once the tokens are in the context window the model treats them identically. The defender's map is what changes. **Direct injection**: the hostile text is in the turn the end user submitted. The adversary is the person using your application — often an authenticated, paying, legitimate account holder — trying to make the model do something *they* are not entitled to do: reveal another tenant's records, invoke a tool above their permission tier, emit content your policy forbids. **Indirect injection**: the hostile text is inside content your application supplied on the user's behalf. A vendor PDF ingested into a retrieval index, a free-text note on a work order, a fetched web page, a tool result, a memory record written during an earlier session. The user may be entirely innocent and may never see the payload. The adversary is whoever can write into that source. ## Why a defender cares Three things differ, and each one shows up in the register. **The entry points differ.** The direct surface is bounded and enumerable — every field of user input, every channel that carries a user turn. You can log all of it in one place. The indirect surface is the transitive set of everyone who can write to anything you ingest. For a plant-maintenance copilot that reads PLC sensor streams, supplier-uploaded equipment manuals and technicians' work-order notes, that set includes people who have never touched your product: a supplier's documentation team, a contractor filing a note, whoever runs a scraped web source. Enumerating writers per source is the actual work of modeling indirect injection, which is why a serious register lists ingestion paths rather than one line saying "prompt injection". **The owner differs.** A direct-injection risk is owned by whoever owns authorization for the application's users. An indirect-injection risk is owned jointly by whoever governs writes into the ingested source — often another team, or an external party with no security relationship to you at all. Assigning that owner is frequently the most useful output of the exercise. **The failure shape differs.** Direct injection is privilege escalation: the user tries to exceed their own authority. Indirect injection is a **confused deputy** — authority and intent come from different parties. The model acts with the caller's or the service's authority on instructions written by a stranger. That framing matters because it tells you the fix cannot be "trust the user less"; it has to be "do not let a path that consumed untrusted content reach a privileged action". ## What each implies for controls For direct injection, the invariant is that a user must not be able to obtain through the model anything they could not obtain by calling your APIs directly. Enforcement lives in the tool layer and the data layer, scoped to the calling identity — never in prompt wording, which is advisory. For indirect injection, the invariant is provenance: the system must know which spans of context came from where, and privileged or irreversible actions must not be reachable from a request that has ingested untrusted spans. The architectural patterns that implement that containment are a separate subject; the threat model's job is to state where the untrusted spans are and what they must not be able to reach. ## The honest 2026 posture Adaptive-attack research — notably the work presented at USENIX Security '26 that broke a dozen published injection and jailbreak defenses at above 90% attack success — has settled the field on treating model-layer defenses as risk reduction rather than as boundaries. Classifiers, delimiters and spotlighting, and instruction-hierarchy training all lower the rate; none of them hold under an attacker who adapts. Direct injection is the harsher case, because the hostile text arrives through the same channel as the legitimate request, so filtering it means filtering your product. Record these controls in the register as mitigations with a residual-risk figure, not as closing the finding. ## Mistakes this split is meant to prevent - Writing one register line called "prompt injection" with a single owner, which collapses two different surfaces and hides every ingestion path. - Treating an internal corpus as trusted because it sits behind the firewall. Trust follows the *writer*, not the storage location or the network zone. - Treating tool output as trusted because you wrote the tool. A tool that returns third-party content returns third-party instructions. - Assuming an output filter covers indirect injection. When the exfiltration channel is a tool call or a rendered link rather than prose, nothing hits the text filter. ## Where the split lands in the standard lists OWASP's GenAI LLM Top 10 2026 keeps prompt injection at LLM01 and describes both forms under it. Once the system has persistent memory and tools, the agentic list re-expresses the same pressure as goal hijack and as memory and context poisoning, because indirect injection through a memory record survives the session that planted it.
- If the retrieval corpus is entirely internal company documents, can you downgrade the indirect-injection risk?Only after you enumerate who can write to it. "Internal" describes network location, not authorship. If suppliers upload manuals, contractors file notes, or a scraper feeds pages in, the corpus carries third-party text and the risk stands. What genuinely downgrades it is a controlled write path with review, plus knowing that the model cannot reach a privileged action from a request that touched that content.
- Does an indirect-injection payload need the user to ever see it?No, and that is what makes it dangerous. The payload only needs to enter the context window on a path the model reads. Retrieval, tool results, memory recall and background summarization all pull content in without rendering it to anyone. Detection therefore cannot depend on user reports; it depends on logging what actually entered context and what actions followed.
- How would you record the two categories differently in a risk register?Direct injection is one entry, owned by the application's authorization owner, with the mitigation stated as per-identity enforcement in the tool and data layers. Indirect injection is a parent entry with one child per ingestion path — corpus, tool, memory store, web fetch — each naming who can write to that source and which privileged actions are reachable downstream of it.
saying these in an interview costs you the question
- Says a strong system prompt prevents both forms
- Calls an internal corpus trusted because it is behind the firewall
- Claims an input classifier closes the finding rather than reducing it
- Treats tool output as trusted because the team wrote the tool
- Thinks indirect injection requires the user to be malicious