skip to content

How do you grade a report that an embedded AI panel's display-name field is injectable, with a screenshot and no reproduction?

level: principalimportance: nice to knowfreq 26%

answer

  1. what did the picture actually show?
  2. demonstrated versus inferred
  3. you cannot re-run in their tenant
  4. reliability is not severity
  5. who owns a seam between two companies?

basics

~20 s

A screenshot of a name inside a panel is an unverified claim about the context path, not a demonstrated exploit. Grade what the evidence licenses, what you can re-run on an account you own, and who owns the seam.

solid answer

~50 s

There are three calls, and they are usually confused for one. The evidentiary call: a screenshot shows a string reaching the screen, which is the delivery path, not a changed task - so ask the reporter for the behaviour change and the timestamps, not more screenshots. The reproduction call: you cannot re-run inside a tenant you do not control, so you rebuild the setup on an account you own; and against a probabilistic system a one-in-five success rate lowers confidence in that particular construction, not in the path, so reliability and severity belong in separate fields. The ownership call is the hard one: in an embedded panel one party renders the field and another assembles the prompt, so the finding lands in a seam neither on-call rota recognises. Deciding who adjudicates it, before arguing about severity, is the part a lead actually owns.

go deeper

for a junior

Know that a screenshot showing your own text in an output is evidence about delivery, not about the model doing something it should not, and that someone has to reproduce it before it is a finding.

for a middle

Be able to say what extra observation you would ask a reporter for, and why re-running inside a customer's live account is not an option available to you.

for a senior

Show you can bound a reproduction rate rather than trade anecdotes, and that you keep observed reliability and reachable impact as separate fields in the report.

for a principal

Own the adjudication: name who arbitrates a seam that sits between the party rendering the field and the party assembling the prompt, and state the written standard for what upgrades an inference into a finding.

## What has actually been reported A report arrives from outside: an outsider set a company display name on their own record, opened a screenshot of a vendor-supplied summary panel inside a host product, and the name is visible in the output. The title says the field is injectable. The report has no reproduction steps, and the account belongs to a tenant nobody on your side controls. Separate three things before grading anything. ## 1. What the evidence licenses The screenshot supports one claim: a string the outsider controlled reached the rendered panel. It does not establish that the value was in the model's context - a panel commonly prints record fields itself - and it does not establish that a directive placed there changes what the assistant does. The claim in the title is two steps ahead of the evidence. What you ask for is therefore not more screenshots but a different kind of observation: output that the surrounding record data cannot account for, whether the effect tracked the reporter's edits across repeats, timestamps against when the panel regenerated, and the field's length. That request is cheap for the reporter and it is the only thing that upgrades an inference into a finding. ## 2. What reproduction can and cannot be Two constraints collide. You cannot legitimately re-run inside a live tenant you do not own, so any verification happens on an account you control with a matching configuration - which is itself a claim you have to check, because an embedded panel is configured per tenant and versioned by the vendor. And the system is probabilistic. A construction that lands once in five attempts is not a one-in-five risk: it is a construction with a low per-attempt rate, in a setting where the attacker retries at no cost and someone else's page view supplies the trigger. So a low rate should reduce confidence in that construction, and should not by itself reduce the severity of what a successful run reaches. Reporting them as separate fields - reachable impact, and observed reliability with a denominator - is what stops the argument from collapsing into a fight over one number. Equally, one failed re-run is not a close. Closing as not-reproducible is honest only when you have run enough attempts on a matching configuration to bound the rate, and when the reporter's account of their setup is complete enough that you know you matched it. Otherwise you have recorded that you did not see it. ## 3. Who adjudicates it This is the call that actually needs a lead, and it is structural. In an embedded multi-tenant panel, one organisation owns the field and its rendering, a second assembles the prompt and ships the panel, and a third - the tenant - owns the data and the account. A finding about a field that was never classified as prompt input by anyone falls precisely between them: the host's team sees a display-name field working as designed, the vendor's team sees text they were handed, and the tenant sees a supplier's record. The predictable failure is that it gets routed as a low-severity display bug because the artefact is a name, and then stalls. So before severity, decide who adjudicates: name a single owner for the seam, and record explicitly that the disagreement is about classification - is a field the builder never counted as prompt input a bug in that field, or a design limit of assembling prompts from tenant data - rather than about whether the screenshot is convincing. ## The standard you should be willing to write down A lead who has been through this once ends up with three written positions, and an interview at this level is largely checking whether you have them: - **What upgrades an inference to a finding** - which observation, not which volume of evidence. - **How reliability is recorded** - a rate with a denominator, separate from impact. - **Who owns a cross-party seam** - decided ahead of the first report, because deciding it during one guarantees a stalled ticket and a frustrated reporter. ## What a weak answer looks like Grading straight off the screenshot, in either direction. Closing it because one re-run failed is as bad as accepting it because the picture is persuasive; both substitute a single observation for a rate, and both leave the ownership question untouched, which is the one that decides whether anything happens at all.

  • The construction reproduces once in five attempts. Does that lower the severity?
    It lowers confidence in that construction, not in the path. Severity should follow what a successful run reaches; reliability belongs in its own field with a denominator. An attacker who can retry at no cost, where somebody else's page view supplies the trigger, does not experience one-in-five as a one-in-five risk.
  • What do you ask the reporter for that does not require touching a live tenant again?
    The behaviour change rather than the echo, the exact field and its length, timestamps relative to the panel regenerating, and whether the effect tracked their edits across repeats. That is enough to rebuild the setup on an account you own, which is where verification should have happened in the first place.
  • When is we cannot reproduce it an honest close?
    When you have run enough attempts on a configuration you have confirmed matches to bound the rate, and the reporter's description of their setup was complete enough to know you matched it. A single failed attempt against a different tenant and a different panel version records only that you did not see it.

saying these in an interview costs you the question

  • Grades the report from the screenshot alone
  • Closes it after one failed re-run
  • Assumes a low reproduction rate means low risk
  • Routes it as a display bug because the artefact is a name
  • Argues severity before anyone owns the seam

context