skip to content

An invoice-intake screen now blocks the wording that drove an unattended payment-hold release, but the same instruction works through the supplier portal - what did that fix buy?

level: seniorimportance: should knowfreq 40%

answer

  1. one path, one wording
  2. a block is a threshold result
  3. silence on a path is not a measurement
  4. the re-test proves carrier independence
  5. the settlement authority never moved

basics

~20 s

It bought coverage of one arrival path against one wording, which is real but narrow. The workflow's authority to commit a settlement write from supplier-supplied text is untouched, so the same instruction re-authored into another supplier carrier still reaches it.

solid answer

~50 s

It bought one thing you can state precisely: on the emailed and EDI document intake, that span now scores above the screen's threshold and is stopped. That is a genuine, measurable reduction, and it is not nothing. What it did not touch is either half of what made the finding expensive - the assistant still reads supplier free text as part of its task, and the workflow still holds authority to release a hold with argument values nobody in the buying company chose. The re-test through the portal comment thread is the proof: the construction is carrier-independent, so covering carriers one at a time buys one integration's worth of coverage per integration while the supplier's cost is one more routine submission. I would report the finding as still open, and be explicit that what changed is the cheapest carrier, not the effect.

code

text · 10 lines
text
finding-2041   unattended settlement workflow released a payment hold
  filed as       prompt injection (LLM01)
  routed to      owner of the assistant's instructions
  change funded  inbound-document screen on the emailed / EDI intake path
  after change   reported wording now scored above threshold on that path
  re-test        [directive span elided], submitted via supplier-portal comment
                 -> same finance write, hold released, 4 of 5 runs
  unchanged      workflow may commit a settlement write with supplier-supplied
                 argument values, unattended, before any batch is skimmed
  status         open; carrier covered, effect not

go deeper

for a junior

Recall that stopping a specific wording on one arrival path is not the same as stopping the behaviour, and that the same sentence can arrive by another route the screen never sees.

for a middle

Explain what a block actually measures - a span scoring above a threshold on the path that inspected it - and why the same span passing elsewhere indicates an uninspected path rather than a weaker one.

for a senior

Demonstrate report discipline: narrow true claims about what changed, a re-test that establishes carrier independence, a run count instead of a bare pass, and a clear statement of why the finding stays open.

for a principal

Set the convention that stops a queue filling with one ticket per carrier, and decide what a partial coverage change may be claimed to have bought when somebody asks for the programme's numbers.

## State what the change actually did, in the narrowest true terms The honest description of a screen added to one intake path is: *on that path, spans scoring above this threshold are stopped, and the reported wording scores above it.* Everything else people want to conclude from a green re-test is an over-read. Get the direction of each claim right, because this is where reports go wrong: - A **block** proves that path scored that span above a threshold. It does not prove the class of construction stopped working, and it does not prove a re-authored variant would also score above it. - A **pass** through the portal proves the span reached the assistant's context by another route and was read as instruction. It does not prove the portal is less protected in some measurable sense - it was never routed through anything, and silence is not a measurement. - The **finance write committing again** proves which call ran, not who chose the argument values. That distinction is the finding. ## What the re-test through a second carrier established Re-authoring the same sentence into the supplier-portal comment thread and getting the same settlement effect establishes something a single reproduction on the original path could not: the construction does not depend on the carrier. The carrier was the part the funded change addressed. Therefore the funded change addressed the part that is interchangeable. That is the sentence to put in the report, and it is not an accusation that the money was wasted. The screen is now a real obstacle on one path, and the class of low-effort attempts that only ever tried that path is now stopped. What the report has to say is that the *finding* is not closed, and why. ## Why the finding stays open The workflow runs overnight and unattended. There is no point between the model reading supplier text and the finance write committing where anything is skimmed. So the boundary on what an obeyed span can achieve is the set of operations the workflow may perform - releasing a payment hold, changing banking details on a supplier record, raising a purchase order against a budget line. None of those moved. The behaviour that made this a finding is still reachable, by a route that costs a supplier one submission. A related trap in triage: **do not re-file the portal reproduction as a new finding.** It is the same defect with a different arrival. Splitting it makes the queue look like a stream of injection reports, each one closed by covering one more carrier, and the count of carriers is set by how many free-text surfaces the business relationship exposes - not by anything the reporting team controls. ## What a re-test can and cannot claim Models are probabilistic. One clean reproduction on the portal path shows the construction worked once against one deployment; it is not by itself a reliability claim. Before reporting, say how many runs were tried and how many landed, and describe what varied between them. A finding that reproduces four times in five reads very differently to an owner than one that landed once, and the difference decides how much of somebody's quarter it can reasonably claim. ## What you say to the person who funded it The engineer's job here is not to win an argument about layers. It is to hand back three things: what the change demonstrably covers, what the re-test demonstrably reached anyway, and which change would make the effect stop. The last one is not the intake team's to make, and saying so is the point of the report rather than a criticism of it. ## The shape of the record A triage record that carries the routing decision, the funded change and the re-test result on one screen is worth more than three tickets, because it shows the label and the mechanism side by side and makes the mismatch visible without argument.

  • The portal reproduction is a different arrival path. Should it be filed as a new finding?
    No. It is the same defect reached by another carrier, and splitting it turns one open finding into a stream of injection tickets, each closed by covering one more free-text surface. The number of those surfaces is set by the business relationship, not by the reporting team, so the queue never converges.
  • The portal re-test landed four times in five. What does that let you claim?
    That the construction reproduces reliably enough to act on, against this deployment, in this window. It is still a statement about one deployment and a probabilistic model, so I would report the run count rather than a bare works or does not work, and say what varied between runs.
  • How would you word the report so it does not read as an attack on the team that funded the screen?
    State what the change demonstrably covers first, then the re-test result, then which change would remove the effect and who owns it. The mismatch is between the label the finding was filed under and the mechanism, not between two teams, and a record showing both side by side makes that obvious without argument.

saying these in an interview costs you the question

  • Calls the finding closed because the reported path is covered
  • Reads a block as proof the class of construction no longer works
  • Re-files each new carrier as a separate injection finding
  • Claims reliability from a single successful reproduction
  • Says the funded change was wasted rather than narrow
  • Treats silence on an uninspected path as evidence of safety

context