skip to content

LLM & GenAI Attack Techniques

This is the attacker's view of LLM applications: injection, jailbreaks, exfiltration, and poisoning, organized around the OWASP LLM Top 10. It is the core craft an AI red-team interview drills into, so interviewers use it to separate people who name attack classes from people who can actually construct and chain them.

on this pageshow

explore

questions

211 · 8 sections

In an email an unattended invoice workflow reads, why is an outright override span the weakest injection?

level: juniorimportance: must knowfreq 72%
basics
~20 s

An outright override announces disobedience, which is the one shape refusal training and screening layers have seen most. Spans that carry supply a reason the task legitimately changed, so the model reclassifies its situation instead of defying its orders.

open as a page

Why doesn't validating a calendar invite's description field stop an assistant from obeying text an outsider wrote there?

level: juniorimportance: must knowfreq 72%
basics
~20 s

The validator checked the field's shape for a screen that only displayed it: length, character set, encoding. Prose passes all of that unchanged, and the assistant now downstream of the same bytes reads prose as something to act on.

open as a page

A logged form value later steers an LLM triage assistant — why did the submission-time scan miss it?

level: juniorimportance: must knowfreq 74%
basics
~20 s

The scan judged the string when nothing could act on it. At submission it was a rejected field value. It became an instruction only when a different component re-read it as operational input it was expected to act on.

open as a page

A team strips HTML from every fetched page — why does an attacker's planted span still arrive intact?

level: juniorimportance: must knowfreq 64%
basics
~20 s

Stripping HTML removes tags, scripts and attributes and keeps the text — and the text is the injection. A sanitiser built to stop a browser executing markup is a no-op against natural-language sentences a model reads as directive.

open as a page

Why do prompt-injection probes target fields like display names and filenames, not the chat box?

level: juniorimportance: must knowfreq 72%
basics
~20 s

The chat box is the surface everyone watches - transcript-logged, sampled and hardened first. A display name or an attachment filename was designed as data, reaches the same model context, and nobody re-reads that path.

open as a page

Why does an attacker's block of hundreds of compliant example exchanges weaken a model's refusal?

level: juniorimportance: must knowfreq 65%
basics
~20 s

A model's refusal is a trained propensity, not an enforced rule. Hundreds of consistent in-context examples showing the model complying supply competing evidence about how this exchange goes, and enough of them outweigh that propensity.

open as a page

In a character-chat product, why can a user-authored persona card get content the model refuses when asked plainly?

level: juniorimportance: must knowfreq 74%
basics
~20 s

A refusal is a learned response to how a request looks, not an enforced rule. A persona card restates the request as a character's line in a scene, a shape the model was mostly trained to continue rather than decline.

open as a page

An attacker mandates a drafting assistant's opening line and bans hedging: why does that change whether it refuses?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A refusal is text the model generates token by token, and it usually begins a characteristic way. Constraining the answer's first words and forbidding qualifiers removes those openings, so the trained tendency to decline has much less to latch onto.

open as a page

Why does a refused request get answered when restated in a low-resource language?

level: juniorimportance: must knowfreq 66%
basics
~20 s

Capability generalises further than safety training does. The model still follows and answers the low-resource form competently, but the alignment data barely covered it, so the refusal behaviour that fires in a widely spoken language never engages.

open as a page

Why does an injected instruction to exfiltrate an unseen email thread need two steps, not one?

level: juniorimportance: must knowfreq 78%
basics
~20 s

The target thread is not in the assistant's context, so the payload cannot simply emit it. It must first make the model look the thread up, then send it. A one-step 'send everything' reaches only what is already loaded — usually a summary.

open as a page

A planted reference reaches an assistant's answer and nobody clicks it — how does data leave?

level: juniorimportance: must knowfreq 66%
basics
~20 s

Displaying is enough. If the answer carries a reference the viewing client resolves on its own — a preview it expands, a resource it embeds — the request goes out while the answer is being drawn. Nobody clicks.

open as a page

In an assistant whose hidden preamble forbids repeating it, why does a translation request still surface the text?

level: juniorimportance: must knowfreq 74%
basics
~20 s

The forbidding line names one operation, repeating, and a translation is a different operation. Nothing marks the preamble as secret; there is only a sentence about one verb sitting beside the text, and the request never uses that verb.

open as a page

Why does an output screen that blocks secrets and raw tool JSON let a review bot describe its own operations in prose?

level: juniorimportance: must knowfreq 62%
basics
~20 s

An output screen matches on shape - credential patterns, key formats, blocks of JSON. A plain-English sentence about which operations a bot has matches none of those, so capability talk leaves as ordinary help text while carrying an inventory.

open as a page

A ticket-triage auto-reply is length-capped per reply - why does that not bound the leak?

level: juniorimportance: must knowfreq 58%
basics
~20 s

A per-reply cap bounds one response, not the total. Each submission is a fresh run over the same hidden context, so many submissions yield many capped slices, reassembled outside the system. Exposure counts per campaign, not per reply.

open as a page

A supplier's invoice text drove an unattended workflow to release a payment hold - what does the prompt-injection label describe?

level: juniorimportance: must knowfreq 68%
basics
~20 s

The prompt-injection label describes only how the instruction arrived: as supplier text the workflow read and treated as an instruction. It says nothing about why that workflow could release a payment hold, which is where the expensive half sits.

open as a page

A user's typed framing got your chat app's hosted model to emit content it normally refuses - which half of that defect can your release fix?

level: juniorimportance: must knowfreq 60%
basics
~20 s

Only the half you ship. Your repository holds the system prompt, the surfaces and what the app does with returned text. The refusal that framing got past is trained behaviour in a supplier's hosted model, with no release of yours.

open as a page

A supplier can re-author one directive sentence into an invoice line, a portal comment or an address block - why is that cheap?

level: middleimportance: should knowfreq 52%
basics
~20 s

The work is the property that makes an assistant read the span as instruction, and that property is channel-independent. The supplier already writes every free-text field of the relationship, so a second carrier costs one more ordinary submission.

open as a page

A filed jailbreak against your hosted chat model reproduces once in five tries - what does that prove?

level: middleimportance: should knowfreq 52%
basics
~20 s

It proves the framing landed once, against one deployment, at one moment. Refusal is a trained propensity sampled at generation time, so a failed retry means that attempt was declined - not that the finding is wrong or the behaviour gone.

open as a page

An invoice-intake screen now blocks the wording that drove an unattended payment-hold release, but the same instruction works through the supplier portal - what did that fix buy?

level: seniorimportance: should knowfreq 40%
basics
~20 s

It bought coverage of one arrival path against one wording, which is real but narrow. The workflow's authority to commit a settlement write from supplier-supplied text is untouched, so the same instruction re-authored into another supplier carrier still reaches it.

open as a page

Why can an attacker exploit a read and a write that each passed an AI assistant's per-capability review?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Because the danger lives in the pair, not either tool. A per-capability review judges each on its own and never sees the combination, and the combination is assembled later, when a deployment wires a private-context read and a network-reaching write into the same session.

open as a page

Why can someone uploading one file to a summarising service outspend its per-request token cap?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A per-request token cap bounds one model call, not one submission. An accepted upload fans out into extraction, a call per unit of content, an aggregation pass and retries, so dozens of individually compliant calls are billed to the operator.

open as a page

Why doesn't ending the session clear an injected fact an email assistant wrote to long-term memory?

level: juniorimportance: must knowfreq 72%
basics
~10 s

Session teardown discards the conversation context, not the durable store. Anything the assistant extracted and saved during that session outlives it by design, so a planted claim is read back into later, unrelated conversations.

open as a page

For an agent's tool calls, what is the difference between an attacker causing a new call and supplying an existing call's arguments?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Causing a new call adds an operation nobody asked for. Supplying arguments leaves the operation exactly as planned and changes only what it acts on. The second is much quieter, because the call itself still looks routine.

open as a page

Why does per-step verification against an agent's recorded plan confirm a goal-substitution attack?

level: middleimportance: must knowfreq 66%
basics
~20 s

Because the check compares each action against the recorded objective, and the recorded objective is what was edited. It tests consistency, not authenticity, so after the edit every action genuinely serves the goal on file and the verifier signs each one off.

open as a page

An attacker's forum post and the vendor's official docs sit in one index — what decides which the retriever returns?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Closeness to the query decides. A first-stage retriever scores each chunk by how near it sits to the question and has no notion of who wrote it, so a forum post and an official page compete on similarity alone.

open as a page

In RAG retrieval, why doesn't a metadata filter reading a document's own status field exclude a planted passage?

level: juniorimportance: must knowfreq 58%
basics
~20 s

A metadata predicate is only as trustworthy as where its value comes from. When status, date or type are copied out of the submitted document, the passage's author sets them too, so narrowing removes honest rivals and keeps the plant.

open as a page

In a wiki assistant that retrieves chunks, why does approving a page's diff not mean a human read what the model reads?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A review and a retrieval consume different artefacts. The reviewer reads an added hunk inside the whole page; the model receives one chunk cut from that page, alone. Approval evidences that a person read the file, not the fragment.

open as a page

A reranking stage is added to a RAG retrieval path — which planted passages does it promote?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A reranking stage drops planted text shaped against an embedding model, which reads as nonsense, and promotes planted text written to read as the ideal answer. It scores relevance to the query, never who wrote the passage.

open as a page

Why does an attacker's encoder-tuned passage rank in first-stage similarity search despite reading badly?

level: juniorimportance: must knowfreq 65%
basics
~20 s

First-stage similarity search embeds query and passage separately and ranks on vector distance alone. That score never reads the text for fluency, authorship or truth, so text fitted to the encoder's geometry can outrank prose written for a human reader.

open as a page

An output harm classifier passed an answer a downstream parser then acted on — what did the screen measure?

level: juniorimportance: must knowfreq 68%
basics
~20 s

It measured harm categories in text: toxicity, violence, self-harm and the like, and found none. The answer was ordinary prose. Its effect came from the parser that gave part of it structural meaning — a property the classifier never scores.

open as a page

Why can a post-generation output screen only truncate a streamed answer an attacker front-loaded?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Streaming sends tokens to the client as they are produced, so a screen that scores the finished answer returns its verdict after the opening has already been delivered. It can stop the remainder; it cannot recall what shipped.

open as a page

An attacker accepts the model's refusal and reads its displayed reasoning — what did the answer-scoped screen never score?

level: juniorimportance: must knowfreq 66%
basics
~20 s

A refusal describes the final answer and nothing else. A screen pointed at the answer never reads the deliberation panel or the trace record beside it, so refused substance and the model's stated objection sit outside its reach.

open as a page

How does an attacker locate a moderation screen's block threshold in a finance assistant that never shows scores?

level: juniorimportance: must knowfreq 68%
basics
~20 s

The score is hidden; the reply is not. A canned block card, a hedged answer and a normal answer are three visible outcomes, so every reply says which band the request landed in, and a few replies bracket the line.

open as a page

How can a request pass a term-based input screen while the generator still answers the original ask?

level: juniorimportance: must knowfreq 72%
basics
~20 s

The screen and the generator read the same string for different purposes. A term-based screen matches surface wording it was trained on; a capable generator resolves paraphrase, referents and framing, so meaning survives a restatement that contains no trained term.

open as a page

Why does text written in non-rendering codepoints still reach a model when no screen shows it?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A screen is a lossy view of stored text. Codepoints such as zero-width format characters and Unicode tag characters produce no glyph, but they are still stored, still copied, and still turned into tokens, so the model reads them.

open as a page

Why does an exact-match term screen miss a blocked word when a user swaps in a look-alike codepoint?

level: juniorimportance: must knowfreq 64%
basics
~20 s

The screen compares codepoints, not appearance. A Cyrillic letter drawn like a Latin one is a different codepoint, so the word no longer equals any stored term and the comparison returns no match, while the glyphs on screen are unchanged.

open as a page

A vendor invoice PDF passed human review — how can extracted text still carry a directive the reviewer never saw?

level: juniorimportance: must knowfreq 68%
basics
~20 s

The reviewer read the document as rendered; extraction deliberately recovers what rendering omits — form-field values, annotation contents, off-canvas or notice-sized runs. Those are two different texts, and only the extracted one reaches the model.

open as a page

Why doesn't a scan of an uploaded call recording catch an injection that appears in its transcript?

level: juniorimportance: must knowfreq 68%
basics
~20 s

The scan inspects audio bytes, and at that moment the instruction text does not exist. Transcription manufactures it afterwards. A clean verdict on the stored file says nothing about the transcript the model later reads.

open as a page

In an assistant that reads pasted screenshots, why is text inside the image untrusted input?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Pasting an image vouches for why the user wanted it, not for who wrote the words in it. A multimodal model reads every legible sentence in the frame, so a planted directive reaches the context as content the model may follow.

open as a page