skip to content

Delivery and Timing

The route from an attacker's keyboard to the model's context: which fields get ingested, what rewrites them in transit, when the span turns directive. Interviewers probe seams nobody called input.

on this pageshow

explore

questions

15

Why doesn't validating a calendar invite's description field stop an assistant from obeying text an outsider wrote there?

level: juniorimportance: must knowfreq 72%

answer

  1. two consumers, one old check
  2. shape, not meaning
  3. the check was sized for a renderer
  4. the renderer could not act; the model can
  5. nobody re-derived the trust call

basics

~20 s

The validator checked the field's shape for a screen that only displayed it: length, character set, encoding. Prose passes all of that unchanged, and the assistant now downstream of the same bytes reads prose as something to act on.

solid answer

~50 s

Validation is always validation *for a consumer*. The invite description was checked years ago for a calendar renderer: a length cap, an allowed character set, maybe some markup stripped so the layout would not break. Those are checks on shape, and they were the right checks, because the worst a sentence in the imperative could do to a display page was be displayed. Adding an email, calendar and contacts assistant put a second consumer downstream of the identical bytes -- one that drafts replies, edits events and rewrites contact records. The consumer changed; the validation did not, because it lives on a write path that predates the assistant and belongs to another team. So "we validate that field" is true and beside the point: it proves the value fit the field, not that it is safe for something that can act.

go deeper

for a junior

Be ready to say what passing a validator actually proves: the value fit the field's shape rules. Recall that the same text is later read by something that can act, and that the check was never written for that reader.

for a middle

An interviewer expects you to trace the field from an outsider's write, through a shape check, into storage, into prompt assembly, and to point at the stage where a trust decision would have had to be re-made -- and note that no such stage exists.

for a senior

Show that you would look for the field's original consumer and its owner before arguing about the model at all, and that you can state the residual plainly: the check is correct for what it was written for and silent about what now reads it.

for a principal

Own the framing that validation is a relation between a field and a consumer, not a property of the field. The organisational consequence is that changing a consumer is a change nobody's process reviews, and that gap crosses two teams.

## The field has two consumers now An email, calendar and contacts assistant reads the owner's diary so it can answer "what is this meeting", draft a reply, or resolve who an attendee is against the address book. To do that, event text -- title, description, location, agenda -- is placed into the model's context. Not all of that text came from the owner. A calendar that accepts meeting requests from outside organisers writes the event, together with whatever prose the organiser typed into its description and location fields, onto the owner's calendar. That write is ordinary. It is what the invite feature is *for*, and no account was compromised to perform it. ## What the validator was written for The description field has been validated for years. When the validator was written, the field had exactly one consumer: a renderer that laid the text out on a screen. A renderer needs the text to be within a length budget, in a character set it can display, and free of markup that would break the layout. Those are the checks that exist. They were never checks on *meaning*, because meaning could not hurt a renderer. So passing validation proves one thing precisely: the value satisfied the predicate the validator encodes. That predicate is about shape -- size, encoding, permitted characters, sometimes a stripped markup subset. Ordinary prose survives every one of them, because prose was never what they were aimed at. ## Why nothing caught the change Three separate facts combine, and none of them is a bug in isolation: 1. **The write path is older than the assistant** and is owned by whoever maintains calendaring, not by whoever shipped the assistant. 2. **Validation happens at write time, for the consumer that existed then.** Nothing about the stored value records what it was validated *against*. 3. **No stage re-derives the trust decision when a new consumer is attached.** Attaching a model downstream is a change in who reads the bytes, and there is no step in anyone's process that asks whether a check written for a passive reader is still the right check for an active one. That third point is the whole of this topic. The defect is not in the validator and not in the model; it is that the question "what is this field trusted for, and by whom?" was answered once and never re-asked. ## Stage by stage | stage | what it sees | what it can decide about meaning | |---|---|---| | invite ingestion | the organiser's raw fields | nothing -- it checks shape | | storage | a stored string, no record of what it was validated for | nothing | | calendar renderer | the text | irrelevant -- it cannot act | | prompt assembly | the same text as context | nothing, unless someone built a stage for it | | the assistant | prose indistinguishable in kind from its own task text | it weighs, it does not enforce | ## What the attacker needs, and what they do not They do not need a bypass, a credential, or a flaw in the validator. They need a field they may already write into by design, which reaches the model's context substantially intact, in front of an assistant whose reach is the owner's reach -- because it acts on the owner's account. The cost is one meeting request. The write leaves the trace of a normal external interaction, because it *is* one. ## The weak answer, and the correction The reflex answer in an interview is "we validate user input." It is the single most common wrong answer here, and it fails on two counts. First, this is not user input in the sense the phrase means -- it was written by an outsider and arrives through a feature designed to let outsiders write it. Second, and more important, validation is not a property a field has; it is a relation between a field and a consumer. The description was validated *as calendar text for a renderer*. Nobody validated it as *instructions for something that can send mail and edit records*, because at the time nothing downstream could do either. ## Saying it precisely "That field is validated" is a statement about a predicate. It is not a trust boundary, it does not travel with the value, and it silently expires the moment something new starts reading the field. A candidate who can state that -- and then say who would have had to notice, and that nobody owns noticing -- has answered the question. A candidate who says "the model should just ignore instructions in data" has answered a different one, and has assumed away the property that made the field eligible in the first place.

  • The team points out the description is capped at two thousand characters. Does the cap help?
    It bounds how much text arrives, not whether the text is directive. A cap constrains volume; the property that matters downstream is that a short, ordinary sentence reads as something to act on. If anything a cap pushes the writer toward concision rather than away from the field. It is a renderer's budget doing exactly its job, which is the point: it was never sized against this consumer.
  • Does it change your assessment if the invite came from an authenticated external organiser?
    Authentication tells you who wrote the text, not that acting on it is safe. In this class the write is legitimate by design -- that is precisely why the carrier is attractive: no credential to steal, no anomaly to detect, and an audit trail that looks like a normal external meeting request. Identity is useful afterwards, for attribution; it does not change what the field is trusted for.
  • Why wouldn't a code review of the assistant have caught this?
    Because the eligible field is not in the assistant's repository. A reviewer of the assistant sees a prompt-assembly step reading calendar text that has been validated on another team's write path for a decade, and there is nothing locally wrong with reading it. The change of consumer is invisible in both diffs -- it exists only in the relation between them.

A mail slot was cut to a size, and letters were checked only for fitting through it. That was enough while the person inside just pinned them to a board -- and stopped being enough when a new resident started doing what they said.

saying these in an interview costs you the question

  • Says the field is safe because it is validated
  • Treats a character-set check as a check on meaning
  • Assumes an outsider cannot write to the owner's calendar
  • Thinks stripping markup removes the sentence
  • Calls it user input, so the user's own problem

context

open as a page

A logged form value later steers an LLM triage assistant — why did the submission-time scan miss it?

level: juniorimportance: must knowfreq 74%

basics

~20 s

The scan judged the string when nothing could act on it. At submission it was a rejected field value. It became an instruction only when a different component re-read it as operational input it was expected to act on.

open as a page

A team strips HTML from every fetched page — why does an attacker's planted span still arrive intact?

level: juniorimportance: must knowfreq 64%

basics

~20 s

Stripping HTML removes tags, scripts and attributes and keeps the text — and the text is the injection. A sanitiser built to stop a browser executing markup is a no-op against natural-language sentences a model reads as directive.

open as a page

Why do prompt-injection probes target fields like display names and filenames, not the chat box?

level: juniorimportance: must knowfreq 72%

basics

~20 s

The chat box is the surface everyone watches - transcript-logged, sampled and hardened first. A display name or an attachment filename was designed as data, reaches the same model context, and nobody re-reads that path.

open as a page

Why is a calendar field outsiders could write into before an assistant existed a better injection carrier?

level: middleimportance: should knowfreq 52%

basics

~20 s

Because its write permission and its validation were both settled for a consumer that could not act. Adding a model downstream changed who reads the field without changing either, and no stage anywhere re-derives that trust decision.

open as a page

An error log line is byte-identical when written and when an LLM triage assistant reads it — what makes it directive only on re-read?

level: middleimportance: should knowfreq 60%

basics

~20 s

Three things change around the unchanged text: the reader now interprets its whole context, the text arrives as the system's own operational evidence rather than a stranger's input, and that reader holds capabilities the write path lacked.

open as a page

Which properties must a span planted on a fetched page have to survive parsing, sanitising and chunking?

level: middleimportance: should knowfreq 44%

basics

~20 s

It must be plain prose, sit in the region a boilerplate remover keeps as main content, depend on no structure or position, and stay whole inside one chunk — properties invariant across a chain nobody outside can observe.

open as a page

A display name an outsider chose appears verbatim in an assistant's summary - what does that prove?

level: middleimportance: should knowfreq 48%

basics

~20 s

An echoed string proves only that the value travelled from the record to the rendered output. The panel may have templated it around the answer, and even if the model read it, echo is not obedience.

open as a page

In an LLM pipeline, what does an attacker gain and give up by deferring a planted span's activation to a later component?

level: seniorimportance: should knowfreq 42%

basics

~20 s

They gain reach: a component with capabilities and framing the arrival path never had, past the boundary where inspection runs. They give up control — no feedback, someone else's schedule, a stage that may not carry the text, no chosen target.

open as a page

A research agent's brief carries a claim planted on a page it fetched — which stage owns the finding?

level: seniorimportance: should knowfreq 37%

basics

~20 s

No stage does. The scraper, sanitiser and converter each met their own specification, so this is not a stage defect: page text chosen at run time entered the context with the same standing as the analyst's request.

open as a page

How do you map which tenant fields reach an embedded AI panel's context without using its chat box?

level: seniorimportance: should knowfreq 40%

basics

~10 s

Vary one field at a time with a distinct harmless marker and read the panel's output for which markers surface and when they stop. Presence is evidence; absence has half a dozen explanations.

open as a page

The team patched the assistant's prompt to ignore invite instructions and closed the finding -- what do you tell them?

level: principalimportance: should knowfreq 40%

basics

~20 s

Say what the patch bought -- lower compliance with the phrasings that were tried, on one build, on one date -- and what it left untouched: who may write the field, what its validation was for, and that nobody owns re-deriving that.

open as a page

Why does an assistant's durable edit to a calendar or contact record rate higher than a wrong answer in one turn?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Because the edit survives the session, carries the owner's authorship to everyone who later reads it, and re-enters the assistant's own context on later turns as the owner's data. A wrong answer stops when the conversation does.

open as a page

An LLM triage run logged an alert-route change — what does that record prove about who chose the arguments?

level: seniorimportance: nice to knowfreq 29%

basics

~20 s

It proves which call ran, with which argument values, under which run — not who chose them. The reason field is model-authored prose, not evidence, and the record cannot distinguish text the system produced from text it quoted.

open as a page

How do you grade a report that an embedded AI panel's display-name field is injectable, with a screenshot and no reproduction?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

A screenshot of a name inside a panel is an unverified claim about the context path, not a demonstrated exploit. Grade what the evidence licenses, what you can re-run on an account you own, and who owns the seam.

open as a page