LLM & GenAI Attack Techniques
This is the attacker's view of LLM applications: injection, jailbreaks, exfiltration, and poisoning, organized around the OWASP LLM Top 10. It is the core craft an AI red-team interview drills into, so interviewers use it to separate people who name attack classes from people who can actually construct and chain them.
on this pageshowhide
explore
- Prompt Injection23 questions
- Authoring the Payload8 questions
- Delivery and Timing15 questions
- Jailbreak Techniques & Taxonomies39 questions
- Anatomy of a Refusal8 questions
- Families That Get Through23 questions
- Half-Life of a Method8 questions
- System-Prompt & Data Exfiltration20 questions
- Making It Speak12 questions
- Egress Channels8 questions
- OWASP LLM Top 108 questions
- Filing the Wrong Owner4 questions
- The Half You Control4 questions
- Agent, Tool-Use & MCP Attacks28 questions
- Turning Text into Effect8 questions
- Composing the Path12 questions
- Memory and Plan8 questions
- RAG & Knowledge-Base Poisoning38 questions
- Getting into the Corpus11 questions
- Being Retrieved19 questions
- Taken as Authority8 questions
- Guardrail & Content-Filter Evasion28 questions
- Input That Passes16 questions
- Tokens Already Sent12 questions
- Multimodal & Encoding-Based Injection27 questions
- Invisible in Plain Sight8 questions
- Not Text at All11 questions
- The Pipeline's Own Edits8 questions
questions
211 · 8 sectionsIn an email an unattended invoice workflow reads, why is an outright override span the weakest injection?
basics
~20 sAn outright override announces disobedience, which is the one shape refusal training and screening layers have seen most. Spans that carry supply a reason the task legitimately changed, so the model reclassifies its situation instead of defying its orders.
Why doesn't validating a calendar invite's description field stop an assistant from obeying text an outsider wrote there?
basics
~20 sThe validator checked the field's shape for a screen that only displayed it: length, character set, encoding. Prose passes all of that unchanged, and the assistant now downstream of the same bytes reads prose as something to act on.
A logged form value later steers an LLM triage assistant — why did the submission-time scan miss it?
basics
~20 sThe scan judged the string when nothing could act on it. At submission it was a rejected field value. It became an instruction only when a different component re-read it as operational input it was expected to act on.
A team strips HTML from every fetched page — why does an attacker's planted span still arrive intact?
basics
~20 sStripping HTML removes tags, scripts and attributes and keeps the text — and the text is the injection. A sanitiser built to stop a browser executing markup is a no-op against natural-language sentences a model reads as directive.
Why do prompt-injection probes target fields like display names and filenames, not the chat box?
basics
~20 sThe chat box is the surface everyone watches - transcript-logged, sampled and hardened first. A display name or an attachment filename was designed as data, reaches the same model context, and nobody re-reads that path.
Why does a per-message content screen pass every turn of a slowly escalating jailbreak?
basics
~20 sA per-message screen scores one message at a time against a threshold, and in a gradual escalation no single message is objectionable. The objectionable object is the whole transcript, which that screen never looks at as one thing.
Why does an attacker's block of hundreds of compliant example exchanges weaken a model's refusal?
basics
~20 sA model's refusal is a trained propensity, not an enforced rule. Hundreds of consistent in-context examples showing the model complying supply competing evidence about how this exchange goes, and enough of them outweigh that propensity.
In a character-chat product, why can a user-authored persona card get content the model refuses when asked plainly?
basics
~20 sA refusal is a learned response to how a request looks, not an enforced rule. A persona card restates the request as a character's line in a scene, a shape the model was mostly trained to continue rather than decline.
An attacker mandates a drafting assistant's opening line and bans hedging: why does that change whether it refuses?
basics
~20 sA refusal is text the model generates token by token, and it usually begins a characteristic way. Constraining the answer's first words and forbidding qualifiers removes those openings, so the trained tendency to decline has much less to latch onto.
Why does a refused request get answered when restated in a low-resource language?
basics
~20 sCapability generalises further than safety training does. The model still follows and answers the low-resource form competently, but the alignment data barely covered it, so the refusal behaviour that fires in a widely spoken language never engages.
Why does an injected instruction to exfiltrate an unseen email thread need two steps, not one?
basics
~20 sThe target thread is not in the assistant's context, so the payload cannot simply emit it. It must first make the model look the thread up, then send it. A one-step 'send everything' reaches only what is already loaded — usually a summary.
A planted reference reaches an assistant's answer and nobody clicks it — how does data leave?
basics
~20 sDisplaying is enough. If the answer carries a reference the viewing client resolves on its own — a preview it expands, a resource it embeds — the request goes out while the answer is being drawn. Nobody clicks.
In an assistant whose hidden preamble forbids repeating it, why does a translation request still surface the text?
basics
~20 sThe forbidding line names one operation, repeating, and a translation is a different operation. Nothing marks the preamble as secret; there is only a sentence about one verb sitting beside the text, and the request never uses that verb.
Why does an output screen that blocks secrets and raw tool JSON let a review bot describe its own operations in prose?
basics
~20 sAn output screen matches on shape - credential patterns, key formats, blocks of JSON. A plain-English sentence about which operations a bot has matches none of those, so capability talk leaves as ordinary help text while carrying an inventory.
A ticket-triage auto-reply is length-capped per reply - why does that not bound the leak?
basics
~20 sA per-reply cap bounds one response, not the total. Each submission is a fresh run over the same hidden context, so many submissions yield many capped slices, reassembled outside the system. Exposure counts per campaign, not per reply.
A supplier's invoice text drove an unattended workflow to release a payment hold - what does the prompt-injection label describe?
basics
~20 sThe prompt-injection label describes only how the instruction arrived: as supplier text the workflow read and treated as an instruction. It says nothing about why that workflow could release a payment hold, which is where the expensive half sits.
A user's typed framing got your chat app's hosted model to emit content it normally refuses - which half of that defect can your release fix?
basics
~20 sOnly the half you ship. Your repository holds the system prompt, the surfaces and what the app does with returned text. The refusal that framing got past is trained behaviour in a supplier's hosted model, with no release of yours.
A supplier can re-author one directive sentence into an invoice line, a portal comment or an address block - why is that cheap?
basics
~20 sThe work is the property that makes an assistant read the span as instruction, and that property is channel-independent. The supplier already writes every free-text field of the relationship, so a second carrier costs one more ordinary submission.
A filed jailbreak against your hosted chat model reproduces once in five tries - what does that prove?
basics
~20 sIt proves the framing landed once, against one deployment, at one moment. Refusal is a trained propensity sampled at generation time, so a failed retry means that attempt was declined - not that the finding is wrong or the behaviour gone.
An invoice-intake screen now blocks the wording that drove an unattended payment-hold release, but the same instruction works through the supplier portal - what did that fix buy?
basics
~20 sIt bought coverage of one arrival path against one wording, which is real but narrow. The workflow's authority to commit a settlement write from supplier-supplied text is untouched, so the same instruction re-authored into another supplier carrier still reaches it.
Why can an attacker exploit a read and a write that each passed an AI assistant's per-capability review?
basics
~20 sBecause the danger lives in the pair, not either tool. A per-capability review judges each on its own and never sees the combination, and the combination is assembled later, when a deployment wires a private-context read and a network-reaching write into the same session.
Why can someone uploading one file to a summarising service outspend its per-request token cap?
basics
~20 sA per-request token cap bounds one model call, not one submission. An accepted upload fans out into extraction, a call per unit of content, an aggregation pass and retries, so dozens of individually compliant calls are billed to the operator.
Why doesn't ending the session clear an injected fact an email assistant wrote to long-term memory?
basics
~10 sSession teardown discards the conversation context, not the durable store. Anything the assistant extracted and saved during that session outlives it by design, so a planted claim is read back into later, unrelated conversations.
For an agent's tool calls, what is the difference between an attacker causing a new call and supplying an existing call's arguments?
basics
~20 sCausing a new call adds an operation nobody asked for. Supplying arguments leaves the operation exactly as planned and changes only what it acts on. The second is much quieter, because the call itself still looks routine.
Why does per-step verification against an agent's recorded plan confirm a goal-substitution attack?
basics
~20 sBecause the check compares each action against the recorded objective, and the recorded objective is what was edited. It tests consistency, not authenticity, so after the edit every action genuinely serves the goal on file and the verifier signs each one off.
An attacker's forum post and the vendor's official docs sit in one index — what decides which the retriever returns?
basics
~20 sCloseness to the query decides. A first-stage retriever scores each chunk by how near it sits to the question and has no notion of who wrote it, so a forum post and an official page compete on similarity alone.
In RAG retrieval, why doesn't a metadata filter reading a document's own status field exclude a planted passage?
basics
~20 sA metadata predicate is only as trustworthy as where its value comes from. When status, date or type are copied out of the submitted document, the passage's author sets them too, so narrowing removes honest rivals and keeps the plant.
In a wiki assistant that retrieves chunks, why does approving a page's diff not mean a human read what the model reads?
basics
~20 sA review and a retrieval consume different artefacts. The reviewer reads an added hunk inside the whole page; the model receives one chunk cut from that page, alone. Approval evidences that a person read the file, not the fragment.
A reranking stage is added to a RAG retrieval path — which planted passages does it promote?
basics
~20 sA reranking stage drops planted text shaped against an embedding model, which reads as nonsense, and promotes planted text written to read as the ideal answer. It scores relevance to the query, never who wrote the passage.
Why does an attacker's encoder-tuned passage rank in first-stage similarity search despite reading badly?
basics
~20 sFirst-stage similarity search embeds query and passage separately and ranks on vector distance alone. That score never reads the text for fluency, authorship or truth, so text fitted to the encoder's geometry can outrank prose written for a human reader.
An output harm classifier passed an answer a downstream parser then acted on — what did the screen measure?
basics
~20 sIt measured harm categories in text: toxicity, violence, self-harm and the like, and found none. The answer was ordinary prose. Its effect came from the parser that gave part of it structural meaning — a property the classifier never scores.
Why can a post-generation output screen only truncate a streamed answer an attacker front-loaded?
basics
~20 sStreaming sends tokens to the client as they are produced, so a screen that scores the finished answer returns its verdict after the opening has already been delivered. It can stop the remainder; it cannot recall what shipped.
An attacker accepts the model's refusal and reads its displayed reasoning — what did the answer-scoped screen never score?
basics
~20 sA refusal describes the final answer and nothing else. A screen pointed at the answer never reads the deliberation panel or the trace record beside it, so refused substance and the model's stated objection sit outside its reach.
How does an attacker locate a moderation screen's block threshold in a finance assistant that never shows scores?
basics
~20 sThe score is hidden; the reply is not. A canned block card, a hedged answer and a normal answer are three visible outcomes, so every reply says which band the request landed in, and a few replies bracket the line.
How can a request pass a term-based input screen while the generator still answers the original ask?
basics
~20 sThe screen and the generator read the same string for different purposes. A term-based screen matches surface wording it was trained on; a capable generator resolves paraphrase, referents and framing, so meaning survives a restatement that contains no trained term.
Why does text written in non-rendering codepoints still reach a model when no screen shows it?
basics
~20 sA screen is a lossy view of stored text. Codepoints such as zero-width format characters and Unicode tag characters produce no glyph, but they are still stored, still copied, and still turned into tokens, so the model reads them.
Why does an exact-match term screen miss a blocked word when a user swaps in a look-alike codepoint?
basics
~20 sThe screen compares codepoints, not appearance. A Cyrillic letter drawn like a Latin one is a different codepoint, so the word no longer equals any stored term and the comparison returns no match, while the glyphs on screen are unchanged.
A vendor invoice PDF passed human review — how can extracted text still carry a directive the reviewer never saw?
basics
~20 sThe reviewer read the document as rendered; extraction deliberately recovers what rendering omits — form-field values, annotation contents, off-canvas or notice-sized runs. Those are two different texts, and only the extracted one reaches the model.
Why doesn't a scan of an uploaded call recording catch an injection that appears in its transcript?
basics
~20 sThe scan inspects audio bytes, and at that moment the instruction text does not exist. Transcription manufactures it afterwards. A clean verdict on the stored file says nothing about the transcript the model later reads.
In an assistant that reads pasted screenshots, why is text inside the image untrusted input?
basics
~20 sPasting an image vouches for why the user wanted it, not for who wrote the words in it. A multimodal model reads every legible sentence in the frame, so a planted directive reaches the context as content the model may follow.