skip to content

Disclosing to a Provider

Some AI findings can only be fixed by a provider you have no contract with, so the report has to help the customer before anyone patches anything. Interviewers ask what you file and what you publish.

on this pageshow

explore

questions

4

During a red-team engagement on a customer application that calls a third-party hosted model endpoint, you reproduce an unsafe response. How do you decide whether the fix owner is the customer or the model provider, and what changes in the report item when the answer is the provider?

level: juniorimportance: must knowfreq 58%

answer

  1. who holds the lever
  2. prompt, retrieval, tools, filters
  3. provider = no date, no verdict
  4. mitigation first, submission second
  5. re-test, no patch event

basics

~20 s

Ask where the behaviour can actually be changed. If the customer's prompt assembly, retrieval data, tool permissions or output filtering could stop it, they own the fix. If only the model's weights or the provider's own safety layer could, the provider owns it, and your report must still give the customer something to apply now.

solid answer

~60 s

Route by **who holds the lever**, not by where the symptom appeared. Walk the request path the customer controls: the system prompt and prompt assembly, whatever text retrieval injects, the permissions on any tool the model can call, and any classifier or regex on input and output. If a change at any of those points removes the effect, it is a customer-side item and gets written like an ordinary application finding with an owner and a date. If nothing the customer controls removes it, the fixer is the provider. That is a different kind of report item, because you have no contract with the provider, no agreed date, and no guarantee they will treat it as a defect at all. So the item carries two things instead of one: the upstream submission (reference, date filed, current status) and a **compensating control the customer can deploy this week** without waiting for anyone. The mitigation is the deliverable; the submission is a status line. Most real cases are mixed: the model is weak, and the customer's design turned that weakness into an impact. Say so, and give both.

go deeper

for a junior

Can say that some findings can only be fixed by the model provider and that the report still has to give the customer a mitigation.

for a middle

Walks the customer-controlled layers one at a time to prove the fix really is not available locally, and writes the mixed case as both items.

for a senior

Designs the compensating control, states the residual risk that remains after it, and sets a re-test because a hosted model can change silently.

for a principal

Sets the firm's standing policy for how provider-owned items are rated, tracked across engagements, and re-tested, so customers are not told to wait indefinitely.

## The four levers the customer actually holds An application built on a hosted model endpoint is a pipeline the customer owns everywhere except one box. Left to right, the owned parts are: **prompt assembly** — the system prompt, the templates, and everything the application concatenates around the user's text before it calls the API; **retrieved context** — whatever a retrieval step pastes in, which becomes attacker-influenced the moment any part of that corpus is user-writable; **tool and permission scope** — the set of functions the model is allowed to invoke and the credentials those functions execute under; and **input/output filtering** — a classifier, a regex, a moderation call, or a schema validator applied on the way in or on the way out. The one box the customer cannot change is the model plus the provider's internal safety stack: the weights, the alignment training, and whatever the provider applies before and after sampling on their side. Routing a finding means testing it against each owned layer before the word "provider" appears anywhere in the write-up. That is the whole discipline, and it is mostly not done. ## What the layer test costs A model finding is a *rate*, not an event, so testing a layer means re-running the same trial count against a modified pipeline. If the effect was established at 12 successes in 50 attempts, an honest check of four layers is four more 50-attempt runs — 200 attempts. Single-turn, that is ~200 API calls; if the effect needs a five- or six-turn build-up, the same 200 attempts is 1,000-1,200 calls, and doubles again if an automated judge scores each transcript. On a mid-tier hosted endpoint the money is rarely the constraint — often well under 50 US dollars. Two other costs bite: **wall clock**, because a 60-requests-per-minute quota turns 1,200 calls into a 20-minute serial run before retries and backoff; and **engineer time**, because each layer test needs a variant of the application stood up with exactly one thing changed. Budget the routing work as its own half-day of the engagement. Teams that do not budget it skip it, assume the provider owns the finding, and file upstream by default. ## Where the number misleads: the clean re-test The dangerous reading is "we added the output filter, re-ran the suite, got 0 of 50 — customer-side fix confirmed." Zero successes in n attempts does not bound the true rate at zero. The standard approximation (the rule of three) puts the top of a 95% interval for zero hits in n trials at about 3/n, so 0/50 is consistent with a real rate as high as roughly 6%. A control that took a 24% rate down to 4% will show zero hits in a 50-attempt run reasonably often, and will usually show one or two, which a tired reader rounds off as noise. Two consequences follow. First, never write "removed" on the strength of a clean small run — write "not observed in 50 attempts; upper bound approximately 6%." Second, and larger: the suite is fixed and does not react, while an attacker does. A filter that cuts your prompt set's rate from 24% to 4% may cut an adaptive attacker's rate not at all, because the attacker rephrases against the filter's observed behaviour. Rate reduction against a fixed suite is evidence about the suite, not about an adversary. ## What changes when the provider is the owner Three properties of a provider-owned item make it a different document. There is **no date**, because you cannot set a remediation deadline for a party who is not in the engagement. There is **no guaranteed verdict**, because the provider may rule the behaviour intended, and their triage is not something you can appeal on the customer's behalf. And there is **no observable patch event**, because a hosted model can change silently with no version you can cite, so "resolved" is something you re-measure, not something you are told. So the item is written mitigation-first: the observed effect and its impact on this application; the evidence that no owned layer removes it; the compensating control the customer can deploy this week (narrow the tool scope, gate the sensitive action behind a human step, drop the capability, or filter the specific effect at the boundary); the residual risk left after that control, with its denominator stated; the upstream submission reference and status as a single line; and a re-test date. ## What to check before committing to "provider owns this" Did you actually build and test each layer variant, or infer it? Is the impact expressed on the customer's asset rather than on the model in the abstract? Is the customer's system prompt merely *asking* for a constraint it could enforce structurally — a very common false attribution upstream? And is the residual-risk number written with what it is a fraction of, so nobody reads "4% of crafted attempts" as "4% of traffic"?

  • The effect only becomes damaging because the application lets the model call a delete operation. Whose item is it?
    The customer's, primarily. The model's weakness is real, but the impact comes from the permission the application granted. Narrow the tool scope; file upstream as a secondary note if you like.
  • What status do you write for an upstream submission that has had no reply?
    Filed on a date, acknowledged or not, no commitment received. State plainly that the customer should not plan around an upstream fix, and point at the compensating control.

Zero failures in fifty tries is like watching a loose stair for one afternoon and seeing nobody trip: it narrows how often the fall happens, but it is not evidence that anyone nailed the board down.

saying these in an interview costs you the question

  • Filing everything upstream and telling the customer to wait for the provider.
  • Putting a remediation deadline on a party who is not in the engagement.
  • Never testing whether a prompt, tool-permission or output-filter change removes the effect.
  • Reporting generic model fallibility with no impact on the customer's application.
  • Treating provider acknowledgement as if it were a fix.

context

open as a page

The provider of a hosted model closes your submission as expected model behaviour rather than a vulnerability, but your client's application still exhibits the effect. What do you do with that item in the client's report?

level: seniorimportance: must knowfreq 52%

basics

~20 s

It stays in the report, re-framed. The item is no longer waiting on a fix; it is a standing property of the platform your client chose. Rate it against the client's application, describe a compensating control the client can apply themselves, and record the residual risk that is left once that control is in place.

open as a page

You are filing a model-behaviour report with the provider of a hosted chat endpoint your client builds on. What evidence does that submission need that a report of a deterministic software bug would not, and why?

level: middleimportance: should knowfreq 44%

basics

~20 s

The behaviour is probabilistic, so one transcript is not a reproduction. Give how many attempts you made and how many succeeded, the exact endpoint and model identifier, the decoding settings and any system prompt, and the date and time window. Without a rate and those conditions, the reader cannot separate your result from noise.

open as a page

Your team wants to publish a writeup of a weakness you found in a third-party hosted model during a client engagement. Classic vulnerability disclosure publishes when a fix ships or an agreed window expires, and here there may be neither a patch you can point at nor a party who owes you a date. How do you set the publication policy?

level: principalimportance: should knowfreq 28%

basics

~20 s

Decide the trigger yourself, in writing, before you file. With no patched version to point at, pick conditions you control: an agreed period since filing, the client's consent as the report's owner, and content that describes the class of weakness and the mitigation without carrying a working prompt anyone can paste.

open as a page