skip to content

OWASP LLM Top 10

You will learn the industry-standard vocabulary that names each LLM risk class and its canonical mitigations, so you can slot a finding into LLM01-LLM10 on the spot. Interviewers lean on it as the shared checklist that structures the rest of the conversation.

on this pageshow

explore

questions

8

A supplier's invoice text drove an unattended workflow to release a payment hold - what does the prompt-injection label describe?

level: juniorimportance: must knowfreq 68%

answer

  1. two halves, two owners
  2. the label names the arrival
  3. the arrival is the cheap half
  4. money moved at the far end
  5. the prompt author cannot unscope a capability

basics

~20 s

The prompt-injection label describes only how the instruction arrived: as supplier text the workflow read and treated as an instruction. It says nothing about why that workflow could release a payment hold, which is where the expensive half sits.

solid answer

~40 s

Injection names the delivery. The supplier wrote directive text into a free-text commercial field on their own submission, the assistant read it as instruction rather than as data to be summarised, and a settlement action followed. That label covers the carrier and the moment of confusion, and nothing else. The other half of the finding is at the far end: an overnight workflow that can release a payment hold with argument values nobody in the buying company chose. The two halves have different owners. The label routes the report to whoever writes the assistant's instructions, while the change that would actually remove the effect belongs to whoever scoped that finance capability. I would describe both halves before I reach for an entry name.

go deeper

for a junior

Be ready to say, in one sentence, that the injection label names how the text got in, and to name separately what the system was then able to do. Practise describing both on a concrete example.

for a middle

An interviewer expects the mechanics: which field the text sat in, why the assistant read it as instruction rather than content, and which call the obeyed output reached. Keep the delivery and the effect verbally separate.

for a senior

Show that you know the label decides who receives the report. Say who can change the arrival, who can change the capability, and why a report naming only the first reaches a team that cannot close it.

for a principal

Own the consequence for the programme: a queue where every LLM finding is labelled by delivery hides where the money actually has to be spent, and the reporting convention is yours to set.

## The two halves of one finding **Prompt injection** is attacker-supplied text that an application feeds its model as *data* being read by the model as *instruction*, so the model does something other than the task the application asked for. It is not the same thing as **jailbreaking**, which aims at the model's trained refusal rather than at the application's own instructions - confusing the two is the most common wrong answer in this whole domain. Injection is called *direct* when the text arrives in the user's own turn and *indirect* when it arrives in content the application retrieves, receives or is handed later. The setting here has no user turn at all. An unattended back-office procurement and settlement workflow takes supplier documents from an intake queue overnight, an assistant reads them, and the workflow drives the finance system directly. There is no chat surface, nobody is watching, and the finance write commits before anyone opens the batch. A finding against that workflow has two halves: 1. **Arrival and confusion.** The supplier put a directive span into a free-text commercial field they legitimately own - a remittance note, a delivery-instruction line, a line-item description. Writing into those fields is not an intrusion; it is the business relationship. The span works because it supplies a plausible business reason for a settlement action, so the assistant treats it as part of the task rather than as content to be reported. 2. **Effect.** The obeyed output reached a privileged operation: a payment hold released, banking details changed on a supplier record, a purchase order raised against a real budget line. Money and an authoritative record, not text. The injection label names half one. ## Why the half you name decides who gets the ticket Triage routes on the label. "Prompt injection" lands with whoever owns the assistant's instructions and the screening on the arrival path they know about. That team can change the app's own wording, and it can tighten a screen on the intake it operates. What it cannot change is whether the workflow is allowed to commit a settlement write whose arguments came out of supplier free text. So a finding filed purely as injection reaches a team that can only address the cheap half. And the halves are priced very differently. The carrier is cheap: one business relationship exposes several free-text fields, a portal comment thread, an address block on the next submission. Re-authoring the same sentence into another of them costs the supplier one more routine submission and no new access. The far end is expensive: changing what the workflow may commit unattended touches the finance close, the batch schedule, and the reason the workflow was built without a human in the first place. ## Getting the direction of each claim right A junior candidate frequently over-reads the evidence. Be careful: - The assistant obeying the span proves the span **reached the model's context and was read as instruction**. It does not prove the finance system was breached, that the supplier holds credentials inside the buying company, or that any access control failed. - A screen blocking the original wording on one intake path proves **that path scored that span above a threshold**. It does not prove the class of construction stopped working. - The workflow committing the write proves **which call ran**, not who chose the argument values - which is precisely the point of the finding. ## When the label really is the whole finding Not every case has a second half. In a consumer chat product with no tools, the only thing that can leave is text; there is no privileged operation at the far end, so the arrival plus the output *is* the finding and the delivery label carries it. The two-halves discipline matters exactly where an obeyed output reaches something that spends money, writes to an authoritative record, or acts without a click. ## What an interviewer is listening for The weak answer stops at "that's prompt injection, entry one". The strong answer separates the arrival from the effect, says who each half routes to, and notes that only one of the two owners can make the behaviour stop. Reciting entry numbers is not the skill being scored; describing the mechanism and naming who can act on it is.

  • Is there a case where the injection label does describe the whole finding?
    Yes, where nothing privileged sits at the far end. In a chat product with no tools, the only thing that can leave is text, so the arrival and the model's output are the entire finding. The two-halves split earns its keep only when an obeyed output reaches a call that spends money, writes to an authoritative record, or acts unattended.
  • The workflow obeyed the supplier's sentence. What does that prove about the finance system?
    Only that the text reached the assistant's context and was read as instruction, and that the workflow already held authority to make that call. It is not evidence of a breach, of stolen credentials, or of an access-control bypass. The supplier used a field they are entitled to write, and the workflow used a capability it was given.
  • Why does this matter more in an unattended batch workflow than in an assistant someone watches?
    Because nothing is skimmed between the model reading the span and the finance write committing. With a person in the loop the effect is bounded by what they let through; overnight it is bounded only by what the workflow is allowed to do, which is exactly the half the delivery label leaves unnamed.

Closing a burglary report with "they came in through the side door" is true and useless: it names the way in and says nothing about why the safe was open.

saying these in an interview costs you the question

  • Treats prompt injection as the root cause rather than the delivery
  • Assumes rewording the assistant's prompt removes the privileged effect
  • Cannot say what the workflow was able to do once it obeyed
  • Calls any unwanted model output prompt injection
  • Confuses injection with getting past a model's trained refusal
  • Reads an obeyed instruction as evidence the finance system was breached

context

open as a page

A user's typed framing got your chat app's hosted model to emit content it normally refuses - which half of that defect can your release fix?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Only the half you ship. Your repository holds the system prompt, the surfaces and what the app does with returned text. The refusal that framing got past is trained behaviour in a supplier's hosted model, with no release of yours.

open as a page

A supplier can re-author one directive sentence into an invoice line, a portal comment or an address block - why is that cheap?

level: middleimportance: should knowfreq 52%

basics

~20 s

The work is the property that makes an assistant read the span as instruction, and that property is channel-independent. The supplier already writes every free-text field of the relationship, so a second carrier costs one more ordinary submission.

open as a page

A filed jailbreak against your hosted chat model reproduces once in five tries - what does that prove?

level: middleimportance: should knowfreq 52%

basics

~20 s

It proves the framing landed once, against one deployment, at one moment. Refusal is a trained propensity sampled at generation time, so a failed retry means that attempt was declined - not that the finding is wrong or the behaviour gone.

open as a page

An invoice-intake screen now blocks the wording that drove an unattended payment-hold release, but the same instruction works through the supplier portal - what did that fix buy?

level: seniorimportance: should knowfreq 40%

basics

~20 s

It bought coverage of one arrival path against one wording, which is real but narrow. The workflow's authority to commit a settlement write from supplier-supplied text is untouched, so the same instruction re-authored into another supplier carrier still reaches it.

open as a page

Triage wants an owner and a fix date for a jailbreak in a supplier's hosted model - what do you commit to?

level: principalimportance: should knowfreq 32%

basics

~20 s

Commit only on the half with a repository: name what your team will ship and when, describe it by what it measurably shifts, and record the model-side behaviour as unresolved rather than scheduled. A supplier's release train takes no date from you.

open as a page

The same jailbreak framing works against other products on the same hosted model - what does that tell you about the defect you filed?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

It locates the defect. The one variable held constant across those products is the model, so the behaviour is model-layer rather than something your prompt or surface created. Your exposure stays yours: the transcript still carries your product's name.

open as a page

A supplier's text drove an unattended payment release; widening the intake screen costs a sprint and removing the workflow's settlement authority costs a quarter - who owns the finding?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Both halves have owners, but only one can remove the effect. File it naming the arrival and the privileged operation separately, route it to the capability owner, and make the expensive option a decision somebody consciously takes.

open as a page