A supplier's invoice text drove an unattended workflow to release a payment hold - what does the prompt-injection label describe?
answer
- two halves, two owners
- the label names the arrival
- the arrival is the cheap half
- money moved at the far end
- the prompt author cannot unscope a capability
basics
~20 sThe prompt-injection label describes only how the instruction arrived: as supplier text the workflow read and treated as an instruction. It says nothing about why that workflow could release a payment hold, which is where the expensive half sits.
solid answer
~40 sInjection names the delivery. The supplier wrote directive text into a free-text commercial field on their own submission, the assistant read it as instruction rather than as data to be summarised, and a settlement action followed. That label covers the carrier and the moment of confusion, and nothing else. The other half of the finding is at the far end: an overnight workflow that can release a payment hold with argument values nobody in the buying company chose. The two halves have different owners. The label routes the report to whoever writes the assistant's instructions, while the change that would actually remove the effect belongs to whoever scoped that finance capability. I would describe both halves before I reach for an entry name.
go deeper
Be ready to say, in one sentence, that the injection label names how the text got in, and to name separately what the system was then able to do. Practise describing both on a concrete example.
An interviewer expects the mechanics: which field the text sat in, why the assistant read it as instruction rather than content, and which call the obeyed output reached. Keep the delivery and the effect verbally separate.
Show that you know the label decides who receives the report. Say who can change the arrival, who can change the capability, and why a report naming only the first reaches a team that cannot close it.
Own the consequence for the programme: a queue where every LLM finding is labelled by delivery hides where the money actually has to be spent, and the reporting convention is yours to set.
## The two halves of one finding **Prompt injection** is attacker-supplied text that an application feeds its model as *data* being read by the model as *instruction*, so the model does something other than the task the application asked for. It is not the same thing as **jailbreaking**, which aims at the model's trained refusal rather than at the application's own instructions - confusing the two is the most common wrong answer in this whole domain. Injection is called *direct* when the text arrives in the user's own turn and *indirect* when it arrives in content the application retrieves, receives or is handed later. The setting here has no user turn at all. An unattended back-office procurement and settlement workflow takes supplier documents from an intake queue overnight, an assistant reads them, and the workflow drives the finance system directly. There is no chat surface, nobody is watching, and the finance write commits before anyone opens the batch. A finding against that workflow has two halves: 1. **Arrival and confusion.** The supplier put a directive span into a free-text commercial field they legitimately own - a remittance note, a delivery-instruction line, a line-item description. Writing into those fields is not an intrusion; it is the business relationship. The span works because it supplies a plausible business reason for a settlement action, so the assistant treats it as part of the task rather than as content to be reported. 2. **Effect.** The obeyed output reached a privileged operation: a payment hold released, banking details changed on a supplier record, a purchase order raised against a real budget line. Money and an authoritative record, not text. The injection label names half one. ## Why the half you name decides who gets the ticket Triage routes on the label. "Prompt injection" lands with whoever owns the assistant's instructions and the screening on the arrival path they know about. That team can change the app's own wording, and it can tighten a screen on the intake it operates. What it cannot change is whether the workflow is allowed to commit a settlement write whose arguments came out of supplier free text. So a finding filed purely as injection reaches a team that can only address the cheap half. And the halves are priced very differently. The carrier is cheap: one business relationship exposes several free-text fields, a portal comment thread, an address block on the next submission. Re-authoring the same sentence into another of them costs the supplier one more routine submission and no new access. The far end is expensive: changing what the workflow may commit unattended touches the finance close, the batch schedule, and the reason the workflow was built without a human in the first place. ## Getting the direction of each claim right A junior candidate frequently over-reads the evidence. Be careful: - The assistant obeying the span proves the span **reached the model's context and was read as instruction**. It does not prove the finance system was breached, that the supplier holds credentials inside the buying company, or that any access control failed. - A screen blocking the original wording on one intake path proves **that path scored that span above a threshold**. It does not prove the class of construction stopped working. - The workflow committing the write proves **which call ran**, not who chose the argument values - which is precisely the point of the finding. ## When the label really is the whole finding Not every case has a second half. In a consumer chat product with no tools, the only thing that can leave is text; there is no privileged operation at the far end, so the arrival plus the output *is* the finding and the delivery label carries it. The two-halves discipline matters exactly where an obeyed output reaches something that spends money, writes to an authoritative record, or acts without a click. ## What an interviewer is listening for The weak answer stops at "that's prompt injection, entry one". The strong answer separates the arrival from the effect, says who each half routes to, and notes that only one of the two owners can make the behaviour stop. Reciting entry numbers is not the skill being scored; describing the mechanism and naming who can act on it is.
- Is there a case where the injection label does describe the whole finding?Yes, where nothing privileged sits at the far end. In a chat product with no tools, the only thing that can leave is text, so the arrival and the model's output are the entire finding. The two-halves split earns its keep only when an obeyed output reaches a call that spends money, writes to an authoritative record, or acts unattended.
- The workflow obeyed the supplier's sentence. What does that prove about the finance system?Only that the text reached the assistant's context and was read as instruction, and that the workflow already held authority to make that call. It is not evidence of a breach, of stolen credentials, or of an access-control bypass. The supplier used a field they are entitled to write, and the workflow used a capability it was given.
- Why does this matter more in an unattended batch workflow than in an assistant someone watches?Because nothing is skimmed between the model reading the span and the finance write committing. With a person in the loop the effect is bounded by what they let through; overnight it is bounded only by what the workflow is allowed to do, which is exactly the half the delivery label leaves unnamed.
Closing a burglary report with "they came in through the side door" is true and useless: it names the way in and says nothing about why the safe was open.
saying these in an interview costs you the question
- Treats prompt injection as the root cause rather than the delivery
- Assumes rewording the assistant's prompt removes the privileged effect
- Cannot say what the workflow was able to do once it obeyed
- Calls any unwanted model output prompt injection
- Confuses injection with getting past a model's trained refusal
- Reads an obeyed instruction as evidence the finance system was breached