skip to content

Prompt Injection

You will learn how untrusted input smuggles instructions into a model's context, both directly in user turns and indirectly through data the app later feeds the model. Interviewers open here because prompt injection is the defining, still-unsolved class of LLM attack and reveals whether you understand the instruction/data confusion at its root.

on this pageshow

explore

questions

23

In an email an unattended invoice workflow reads, why is an outright override span the weakest injection?

level: juniorimportance: must knowfreq 72%

answer

  1. command versus premise
  2. the loudest shape is the trained one
  3. it never mentions the earlier rules
  4. a reason the task changed, not a demand

basics

~20 s

An outright override announces disobedience, which is the one shape refusal training and screening layers have seen most. Spans that carry supply a reason the task legitimately changed, so the model reclassifies its situation instead of defying its orders.

solid answer

~50 s

An override span states that the model should stop doing what it was told, and that shape loses three ways: it is the densest region of refusal training data, it is the densest region of the corpus an input screen was built from, and it forces the model to notice a conflict and pick a side — a comparison the application's own instructions usually win. A span that carries does the opposite. It never mentions the earlier orders. It presents as ordinary content of the kind the run was built to read as fact, such as a counterparty asserting that an order was amended, and supplies a reason the task in front of the model is now a different task. The model is not disobeying; it has reclassified its situation. That is the difference between a command and a premise.

go deeper

for a junior

Be ready to state the contrast cleanly: a span that commands the model to abandon its orders versus a span that supplies a reason the job changed, and why the second is the one that carries.

for a middle

Expect to explain the three separate failure modes of the override shape — trained resistance, screening-corpus density, and forcing an explicit conflict — rather than lumping them into 'models are trained against it'.

for a senior

Show that you can say where the construction stops working: it needs a reason plausible inside the corpus the run reads, and a task that admits an alternative coherent reading. Name the obstacle it got past without proposing the fix.

for a principal

Own the reading of the evidence. One successful run is a data point about one deployment, and a span that was acted on tells you it reached the context and cohered — not that anything upstream was missing.

## The setting An unattended supplier-invoice and purchase-order workflow reads inbound email. Every run opens by restating the buyer's standing routing and approval rules, then states the run's own task: extract the amount, match it to an open purchase order, file a verdict. Nobody reads the correspondence on the routine path; the filed verdict is the output, and later readers treat it as already checked because a machine signed it. Anything an attacker writes into an email body or signature block on an open thread arrives as data the run was built to read as fact. ## Two shapes an injected span can take **Shape one — the override.** Text addressed to the model, referring to its instructions, telling it to stop following them. This is the shape almost every candidate names first. **Shape two — the premise.** Text that never mentions the model or its instructions at all, that presents as ordinary content of the kind the run consumes, and that supplies a reason the task in front of the model is a different task from the one it assumed — a counterparty stating that the order was amended and that a different line now applies. ## Why the override shape loses Three things work against it, and they fail at different stages, so keep them apart. **It is the densest region of refusal training data.** That phrasing family has been public for years and is exactly what instruction-following adversarial post-training aims at. Models resist it better than almost anything else in this domain. **It is the densest region of screening data.** An input screen, whether it matches patterns or scores with a classifier, was built from a corpus of this shape, because that is the corpus that exists. A span in this family is the easiest thing such a layer has to score. **It poses the conflict out loud.** It requires the model to notice that two sets of instructions disagree and to choose between them. The application's instructions are the ones a model prefers in that comparison, and they were supplied in the turn it treats as most authoritative. Forcing an explicit choice is asking to lose the comparison you most reliably lose. ## What a carrying span supplies instead A span that carries never enters that comparison. It supplies three things. **A source worth crediting.** It arrives on a thread the workflow already corresponds with, from a party the run has reason to read, and it reads like that party. Nothing about it is anomalous to a model reading correspondence. **A reason the task changed.** Not a demand that the rules be dropped — an assertion about the world from which a different task follows. The standing rules remain true; the situation they govern is claimed to be different. **A job that is coherent on its own.** The model can carry the alternative reading through to an output that satisfies the run's stated task. A reading that leaves the model unable to finish gets discarded in favour of the original one. Notice what is absent: the span never contradicts the standing rules out loud. Compliance and obedience are made to coincide, which is why there is no disobedience for a trained preference to resist. ## Where the construction stops working It is target-specific. The reason has to be plausible inside the corpus the run actually reads, which means it depends on the buyer's vocabulary, the state of the thread, and what the workflow is for. That is reconnaissance cost, not model-trick cost, and it does not port to the next buyer. It also depends on the run's task admitting an alternative reading. If the job is narrow enough that no other coherent task exists — return one field from one document and nothing else is an output — a premise has nothing to attach to. And it says nothing about position or precedence. Where the span sits governs whether it survives to be read at all, which is a different question with a different answer. What decides adherence here is semantic: whether the model finds the alternative frame more coherent than the original one. ## Reading the result correctly A span the model acted on proves the span reached the context and was found coherent. It does not prove that a screening layer was absent, that the mailbox or the store was compromised, or that the model's instruction ordering was inverted. And a construction that worked once against one deployment has worked once against one deployment — with a probabilistic model, that is a data point, not a reliability claim.

  • Does the model have to believe the claim for this to work?
    Belief is the wrong frame. The model weighs readings of the material in front of it, and the span only has to make the alternative reading the more coherent one — a claim consistent with the thread, the vocabulary and the run's stated job. Nothing verifies it, because the run's task is to read correspondence as fact, so plausibility inside that corpus is the whole bar.
  • What if the claimed reason is checkable and demonstrably false?
    It only matters if something in the pipeline checks it. An unattended run that derives its verdict from the correspondence it was handed has no second source to contradict the claim, and the model has no notion of who wrote a passage or whether they were entitled to say it. Checkability becomes relevant only once a later reader re-derives the claim, which on the routine path nobody does.
  • So does bare override phrasing never work?
    It works where nothing is in the way — a deployment with no input screen and a model that happens to comply. It is the cheapest thing to try, which is why it is tried first, and its failure is informative: it tells you something resisted, though not whether the model refused or a screening layer blocked. Those are different events and telling them apart is a separate measurement.

A forged order tells the clerk to break the rule. A forged fact tells the clerk the rule now applies to a different case — and never asks them to break anything.

saying these in an interview costs you the question

  • Says the payload must tell the model to ignore previous instructions
  • Treats it as a precedence problem the system prompt can settle
  • Assumes anything a model obeyed must have looked like an instruction
  • Thinks a longer or more emphatic command performs better
  • Believes the span has to contradict the standing rules to beat them

context

open as a page

Why doesn't validating a calendar invite's description field stop an assistant from obeying text an outsider wrote there?

level: juniorimportance: must knowfreq 72%

basics

~20 s

The validator checked the field's shape for a screen that only displayed it: length, character set, encoding. Prose passes all of that unchanged, and the assistant now downstream of the same bytes reads prose as something to act on.

open as a page

A logged form value later steers an LLM triage assistant — why did the submission-time scan miss it?

level: juniorimportance: must knowfreq 74%

basics

~20 s

The scan judged the string when nothing could act on it. At submission it was a rejected field value. It became an instruction only when a different component re-read it as operational input it was expected to act on.

open as a page

A team strips HTML from every fetched page — why does an attacker's planted span still arrive intact?

level: juniorimportance: must knowfreq 64%

basics

~20 s

Stripping HTML removes tags, scripts and attributes and keeps the text — and the text is the injection. A sanitiser built to stop a browser executing markup is a no-op against natural-language sentences a model reads as directive.

open as a page

Why do prompt-injection probes target fields like display names and filenames, not the chat box?

level: juniorimportance: must knowfreq 72%

basics

~20 s

The chat box is the surface everyone watches - transcript-logged, sampled and hardened first. A display name or an attachment filename was designed as data, reaches the same model context, and nobody re-reads that path.

open as a page

Why does an attacker's span placement differ between a whole-file summariser and a chunk retriever?

level: middleimportance: must knowfreq 62%

basics

~20 s

The unit of survival changes. A summariser forwards the file minus whatever truncation drops, so one occurrence anywhere inside the surviving window is enough. A retriever forwards a single chunk, so the span must sit whole inside one cut region.

open as a page

Why is page-one placement of an injected span in an uploaded document not a general rule?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Page one matters only if the path reading the file starts there. A chunk retriever forwards one selected passage, not the opening page; a summariser over a long file keeps only what fits its budget. Placement follows the path.

open as a page

An invoice workflow restates its rules each run — why doesn't a later position give an injected span leverage?

level: middleimportance: should knowfreq 46%

basics

~20 s

Position is not what decides adherence. Recency is weak and length-dependent; leverage comes from what a span claims — a creditable source, a reason the task changed, a finishable job — not from where it sits.

open as a page

Why is a calendar field outsiders could write into before an assistant existed a better injection carrier?

level: middleimportance: should knowfreq 52%

basics

~20 s

Because its write permission and its validation were both settled for a consumer that could not act. Adding a model downstream changed who reads the field without changing either, and no stage anywhere re-derives that trust decision.

open as a page

An error log line is byte-identical when written and when an LLM triage assistant reads it — what makes it directive only on re-read?

level: middleimportance: should knowfreq 60%

basics

~20 s

Three things change around the unchanged text: the reader now interprets its whole context, the text arrives as the system's own operational evidence rather than a stranger's input, and that reader holds capabilities the write path lacked.

open as a page

Which properties must a span planted on a fetched page have to survive parsing, sanitising and chunking?

level: middleimportance: should knowfreq 44%

basics

~20 s

It must be plain prose, sit in the region a boilerplate remover keeps as main content, depend on no structure or position, and stay whole inside one chunk — properties invariant across a chain nobody outside can observe.

open as a page

A display name an outsider chose appears verbatim in an assistant's summary - what does that prove?

level: middleimportance: should knowfreq 48%

basics

~20 s

An echoed string proves only that the value travelled from the record to the rendered output. The panel may have templated it around the answer, and even if the model read it, echo is not obedience.

open as a page

An injected supplier email took ten drafts, mostly learning the buyer's own routing vocabulary — what follows?

level: seniorimportance: should knowfreq 38%

basics

~20 s

The cost sat in reconnaissance, not in model behaviour, so the construction is target-specific rather than model-specific. It is unlikely to port to another buyer, more likely to survive a model change, and by design it reads as ordinary correspondence.

open as a page

A filed injection finding reproduces once in five runs of the same uploaded file — is it a finding?

level: seniorimportance: should knowfreq 42%

basics

~20 s

It can be, but not yet. First establish whether the span reached the model on the runs that worked and not on the others. Intermittent arrival is a plumbing result; intermittent obedience after arrival is a different result with a different owner.

open as a page

In an LLM pipeline, what does an attacker gain and give up by deferring a planted span's activation to a later component?

level: seniorimportance: should knowfreq 42%

basics

~20 s

They gain reach: a component with capabilities and framing the arrival path never had, past the boundary where inspection runs. They give up control — no feedback, someone else's schedule, a stage that may not carry the text, no chosen target.

open as a page

A research agent's brief carries a claim planted on a page it fetched — which stage owns the finding?

level: seniorimportance: should knowfreq 37%

basics

~20 s

No stage does. The scraper, sanitiser and converter each met their own specification, so this is not a stage defect: page text chosen at run time entered the context with the same standing as the analyst's request.

open as a page

How do you map which tenant fields reach an embedded AI panel's context without using its chat box?

level: seniorimportance: should knowfreq 40%

basics

~10 s

Vary one field at a time with a distinct harmless marker and read the panel's output for which markers surface and when they stop. Presence is evidence; absence has half a dozen explanations.

open as a page

The team patched the assistant's prompt to ignore invite instructions and closed the finding -- what do you tell them?

level: principalimportance: should knowfreq 40%

basics

~20 s

Say what the patch bought -- lower compliance with the phrasings that were tried, on one build, on one date -- and what it left untouched: who may write the field, what its validation was for, and that nobody owns re-deriving that.

open as a page

In a document you authored, you cannot see where the splitter cuts — what follows for placing a span?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Placement becomes probabilistic rather than exact, so the construction shifts to several short, self-contained occurrences spread across the file. Each copy raises the chance one lands whole inside a cut region, and each adds surface.

open as a page

Why does an assistant's durable edit to a calendar or contact record rate higher than a wrong answer in one turn?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Because the edit survives the session, carries the owner's authorship to everyone who later reads it, and re-enters the assistant's own context on later turns as the owner's data. A wrong answer stops when the conversation does.

open as a page

An LLM triage run logged an alert-route change — what does that record prove about who chose the arguments?

level: seniorimportance: nice to knowfreq 29%

basics

~20 s

It proves which call ran, with which argument values, under which run — not who chose them. The reason field is model-authored prose, not evidence, and the record cannot distinguish text the system produced from text it quoted.

open as a page

An invoice workflow's owner will add a standing rule forbidding direction from supplier email — what do you tell them?

level: principalimportance: nice to knowfreq 27%

basics

~20 s

Tell the owner what the rule buys: spans shaped as commands get costlier. It does nothing about spans shaped as facts, since reading correspondence as fact is the run's job. Then settle who owns the residual.

open as a page

How do you grade a report that an embedded AI panel's display-name field is injectable, with a screenshot and no reproduction?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

A screenshot of a name inside a panel is an unverified claim about the context path, not a demonstrated exploit. Grade what the evidence licenses, what you can re-run on an account you own, and who owns the seam.

open as a page