A logged form value later steers an LLM triage assistant — why did the submission-time scan miss it?
answer
- two moments, not one
- nothing could act on it yet
- the scan ran, then stopped running
- same bytes, different reader
- the log line lost who typed it
basics
~20 sThe scan judged the string when nothing could act on it. At submission it was a rejected field value. It became an instruction only when a different component re-read it as operational input it was expected to act on.
solid answer
~50 sA screen at the submission boundary sees text in the only framing that exists at that instant: a form field that failed validation and is about to be discarded. Nothing downstream of that boundary can execute anything, so there is no instruction there to catch — only prose that reads oddly. The value is then echoed verbatim into the application's error log by code whose only job was formatting a message. Hours later a triage assistant queries that log store and reads the same characters as operational evidence it is supposed to act on. The activation is entirely on the second read: same bytes, different reader, different capabilities. This is what makes "we scan everything on ingestion" the wrong answer — the scan ran where the span was not yet an instruction to anybody, and it cannot see the framing the second system will supply.
go deeper
Be ready to say plainly that the string was not an instruction when it was scanned, and that a second component reading it later is what made it one. Do not stop at the phrase indirect injection.
An interviewer expects you to trace the actual path — submission, formatting into a log message, later query into a model's context — and to say what each stage could and could not know about the text it handled.
Show that you can name the seam rather than blame a component. Say what the submission screen legitimately established, and why nothing between it and the re-read was in a position to carry provenance forward.
Be able to frame this as a property of the pipeline's shape rather than a bug in one service, and to say what class of change actually moves it versus what only relabels the same seam.
## Two moments, not one A second-order (deferred) injection deliberately separates the moment a span **arrives** from the moment it **acts**. The attacker writes into a field that is validated, rejected and logged — a display name, a malformed request parameter, a header value. Nothing about that write is an attack yet. The attack is that some later component, which the attacker never touches and often does not know the shape of, re-reads the stored text in a setting where text is treated as something to follow. Walk the path in an incident-triage deployment over an observability stack: 1. **Submission.** A form value fails validation. An input screen at that boundary scores the text and lets it through, or never sees it at all because rejected values are not what that screen was pointed at. 2. **Storage.** The application logs the failure. The logging call is ordinary string formatting; it embeds the submitted characters in a message so an engineer can later see what was rejected. Nothing in that code makes a trust decision, because in the world it was written for, nothing downstream of a log line could act. 3. **Re-read.** A triage assistant queries the log store, pulls recent error strings into its context alongside alert payloads, and produces a shift summary — and on the low-severity path, adjusts alert routing on its own. The span was inert at steps 1 and 2 and directive at step 3. ## What the submission scan actually judged This is the point candidates get backwards. A screen at the submission boundary establishes exactly one thing: **this text scored below a threshold, at that boundary, against whatever that screen was built to catch.** It does not establish that the text is harmless, and it certainly does not establish that no later component will read it as directive. Passing a screen is a measurement, not a property of the string. And the screen was not being careless. At submission there is no tool to call, no alert route to change, no summary a human will act on. The text is not an instruction because there is no instruction-follower in the room. You cannot ask a scan to catch an instruction that does not exist yet. ## Why the delay is the whole technique The attacker is not evading the scan by disguising the text. They are evading it by **choosing a moment when it is not running**. The screen ran once, at submission. It does not run again when the log store is queried, because a query against your own logs was never modelled as untrusted input — that data is, by everyone's mental model, something the system produced about itself. The span also picks up something on the way: **framing**. By the time the assistant reads it, the characters are no longer labelled "a value a stranger typed into a form". They are a line in the application's own error log, retrieved as operational evidence. Provenance was not stripped maliciously; it was simply never carried, because the formatting code had no reason to carry it. ## Where this class stops working It is not universal, and an interviewer will probe the limits: - **The span has to survive the trip.** If the intermediate stage paraphrases, truncates or summarises rather than quoting, the characters may not reach the second reader intact. - **The second reader has to have somewhere to go.** If the component that re-reads the log can only print text to a human who is skimming, the ceiling is a misleading sentence in a summary rather than an action. - **Timing belongs to somebody else.** Activation happens when the overnight summary regenerates or when an engineer opens a console — on a schedule the attacker does not control and cannot observe. ## The claim to state carefully A model acting on a retrieved log line proves the line reached the model's context. It does not prove the log store was breached, and it does not prove the submission screen failed at its own job. It proves the pipeline has a boundary where text that entered as data is read back as evidence, with no step in between that re-derives who wrote it. That sentence — not a list of injection families — is what a first-round answer should land on.
- The screen at submission did its job correctly. So where would you say the defect lives?At the seam, not at either end. The submission screen judged text in a setting where nothing could act on it, and the log query reads text in a setting where something can — but no step between them re-derives who wrote the characters. Both components are individually reasonable; the pipeline never asks the question a second time.
- Does it matter that the field was rejected rather than accepted?It helps the attacker. A rejected value skips the validation and storage paths a valid value would take, and it lands in an error log — a store nobody treats as user-authored content. The rejection is what routes the characters into the carrier that the later component reads as the system's own account of itself.
- Is this the same thing as an attacker typing the instruction into the chat box?No. A direct injection arrives in the user's own turn, aimed at the application's instructions in that same exchange, and it is delivered on the channel most likely to be watched. Here the attacker never speaks to the assistant at all; the text arrives on a channel nobody classed as input, and fires against whoever consumes the derived artefact next.
A note slipped into a filing cabinet is not an order. It becomes one when somebody who takes orders from that cabinet reads it years later.
saying these in an interview costs you the question
- Says ingestion scanning covers it because every field is scanned
- Claims passing a screen proves the text is harmless
- Treats logs and retrieved evidence as trusted because the system wrote them
- Assumes the attacker must reach the chat interface to inject anything
- Confuses this with the model refusing a request it was trained to decline