skip to content

Finding the Untrusted Seam

The chat box has the most eyes on it; the productive surfaces are a filename, a display name, a ticket title. Interviewers use it to see whether you can read an app you cannot see inside.

on this pageshow

explore

questions

4

Why do prompt-injection probes target fields like display names and filenames, not the chat box?

level: juniorimportance: must knowfreq 72%

answer

  1. which surface has the most eyes on it
  2. the builder wrote their wording against one box
  3. a filename is text too
  4. transcripts cover the typed path only

basics

~20 s

The chat box is the surface everyone watches - transcript-logged, sampled and hardened first. A display name or an attachment filename was designed as data, reaches the same model context, and nobody re-reads that path.

solid answer

~50 s

Both surfaces put attacker-controlled text into the same context, but they are not equally productive. The chat box is the one the builder had in mind: the application's own wording was written against it, typed turns land in a retained transcript that gets sampled per tenant, and anything found there is usually already on someone's list. A tenant-supplied company display name, a note field, a locale string or an attachment filename that gets rendered into the prompt from a database row is the same untrusted text arriving on a path nobody classified as prompt input, so there is no transcript of it and no per-turn review. The productive question is therefore not `is this assistant injectable` but `which fields does the builder not think of as prompt input`. The cost is feedback: a record write only produces an observable result when somebody opens the panel.

go deeper

for a junior

Be ready to name three fields on a real screen that plausibly get rendered into an assistant's prompt, and to say why text pulled from a record is no safer than text somebody typed.

for a middle

Explain the mechanics of the difference: the typed turn becomes a reviewed transcript entry, the record value becomes a row that some other user's session later reads. Same context, different observation.

for a senior

Show that you price the surfaces rather than listing them. Say which one you would spend probes on given a slow feedback loop, and what your result would describe - one tenant, one panel version.

for a principal

Own the framing that this is a classification problem, not a filtering one: the field was never counted as prompt input by anyone in the org chart, which is why nothing covers it and why nobody's on-call recognises the finding as theirs.

## The setting Take a vendor-supplied *summarise this account* panel that ships inside somebody else's CRM or marketplace. When a user opens a record, the panel assembles a prompt out of that record: the company display name, a free-text note field, a locale or region string, the filenames of the attachments, some row counts, and the vendor's own framing wording. An outside party who appears on that record - a supplier, an applicant, a marketplace counterparty - can set several of those values on their own side, and will never see the vendor's prompt assembly. ## What a seam is A seam, for this purpose, is any field whose value someone outside the trust relationship can set and whose text ends up rendered into the model's context. The chat box is one seam. So is every field listed above. The chat box is the seam the whole industry looks at, which is exactly what makes it the least interesting one to spend probes on. ## Why the chat box is the worst place to probe None of the reasons is filtering. They are all about attention. 1. **The wording was written against it.** Whatever the builder wrote to anticipate a hostile instruction, they wrote it imagining a typed turn. 2. **It is observed.** A typed turn becomes a transcript entry: retained, sampled, reviewed per tenant. A value written onto a record does not become a transcript; it becomes a row. 3. **It is fixed first.** Anything demonstrable there is usually already known, so the marginal value of a finding is low. 4. **It is bound to your session.** The typed turn is attributable to whoever typed it; a value sitting on a record is read whenever some other person opens the panel, and the read belongs to them. ## Why the unglamorous fields are productive The fields were designed for a consumer that could only display them, so nobody in the building has ever counted them as part of *the prompt*. That mental model, not a missing control, is what leaves the path unwatched: the review that covers the chat path stops at the boundary of the chat path. The reviewer who samples transcripts never sees a record write, and the person who approves a record write is not thinking about a model. The standard vocabulary calls the typed case direct and the assembled case indirect or second-order; OWASP's Top 10 for LLM Applications catalogues the class as LLM01 Prompt Injection. Assume that vocabulary - the interesting part is which concrete fields on a real screen turn out to be in the assembled set. ## What it costs the person doing it Probing an unwatched field is slower, not free. - **Latency.** You get no per-turn loop. A written value only produces an observable result when a human on the other side opens the panel, which may be hours or never. - **Durability.** The value sits in someone's live account until it is changed. Unreviewed is not the same as untraceable: application logs may capture the assembled prompt even when no person reads them. - **Locality.** What you learn describes one tenant's configuration at one moment. A different tenant may have a different field set; the vendor may ship a new panel next week. ## The answer that fails The weak answer is *you try prompt injection on the chat box*. It is not wrong that the chat box accepts injection; it is wrong about where the work is. Someone who says it has usually only ever driven a chat UI, and their next sentence tends to be that the fix is better wording in the system prompt. ## What a good answer sounds like Name two or three concrete fields on a screen you have actually seen, say why each one is plausibly in the assembled prompt, and say what the difference in observation between the typed path and the field path buys the person probing. Being able to say *I do not know whether the filename is in there, and here is the cheapest way to find out* is worth more at this level than a list of injection families.

  • Does landing in an unreviewed field mean the probe leaves no trace?
    No. The value is a durable row in a live account, visible to anyone who opens the record, and it stays there until changed. Unreviewed is not unlogged either - the application may retain the assembled prompt or the panel output even though no human samples that path. The claim you can defend is that nobody is currently looking, not that there is nothing to look at.
  • What does the chat box still tell you that a field probe cannot?
    It gives a fast loop. You learn the assistant's surface behaviour - how it phrases things, how it declines - in seconds rather than waiting on somebody to open a record. What it does not give you is transferable results: a span that is accepted in a typed turn may never survive the assembly path a record field takes, and the reviewed transcript makes that learning expensive.

Testing the front door under the camera tells you the door is a door. The delivery hatch nobody has looked at since it was installed is where the interesting answer is.

saying these in an interview costs you the question

  • Says the chat box is the main injection surface
  • Treats a database field as trusted because a form validated it
  • Assumes an unreviewed path is an untraceable one
  • Confuses aiming at the application's instructions with aiming at the model's refusal training
  • Cannot name a single concrete field other than the chat input

context

open as a page

A display name an outsider chose appears verbatim in an assistant's summary - what does that prove?

level: middleimportance: should knowfreq 48%

basics

~20 s

An echoed string proves only that the value travelled from the record to the rendered output. The panel may have templated it around the answer, and even if the model read it, echo is not obedience.

open as a page

How do you map which tenant fields reach an embedded AI panel's context without using its chat box?

level: seniorimportance: should knowfreq 40%

basics

~10 s

Vary one field at a time with a distinct harmless marker and read the panel's output for which markers surface and when they stop. Presence is evidence; absence has half a dozen explanations.

open as a page

How do you grade a report that an embedded AI panel's display-name field is injectable, with a screenshot and no reproduction?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

A screenshot of a name inside a panel is an unverified claim about the context path, not a demonstrated exploit. Grade what the evidence licenses, what you can re-run on an account you own, and who owns the seam.

open as a page