skip to content

Your team wants to publish a writeup of a weakness you found in a third-party hosted model during a client engagement. Classic vulnerability disclosure publishes when a fix ships or an agreed window expires, and here there may be neither a patch you can point at nor a party who owes you a date. How do you set the publication policy?

level: principalimportance: should knowfreq 28%

answer

  1. no patch event, no timer
  2. triggers you control
  3. client consents, client anonymised
  4. class not instance
  5. cannot claim a fix
  6. named sign-off, decided early

basics

~20 s

Decide the trigger yourself, in writing, before you file. With no patched version to point at, pick conditions you control: an agreed period since filing, the client's consent as the report's owner, and content that describes the class of weakness and the mitigation without carrying a working prompt anyone can paste.

solid answer

~1 min

Three questions, settled up front. **What triggers publication.** You cannot key off a fix event, because a hosted model may change silently or never. So the firm's policy has to name triggers it owns: a fixed period after filing, an explicit decline or non-response, or the moment the behaviour stops reproducing. Say which one fired, and say that you could not verify a fix — do not claim one. **Who consents.** The findings came from a client engagement, so the client owns the report and any detail that identifies their application, their data or their configuration. Publication needs their written agreement, and usually means stripping the engagement down to the model behaviour alone. **What the artefact contains.** Publish the class of weakness, the conditions under which it appeared, the rate if you measured one, and the mitigations readers can apply. Do not publish a copy-paste payload: unlike a software bug where the fix precedes the writeup, here there is often no fix, so the string stays live for everyone downstream. And expect the honest outcome that the provider may consider it intended behaviour. Publishing then is a claim about a design property, not about a vulnerability, and the wording should say so.

go deeper

for a junior

Knows publication needs the client's permission and should not include a copy-paste attack string.

for a middle

Can name a trigger and anonymise the engagement, and understands that no observable patch event exists to key off.

for a senior

Chooses triggers the firm controls, phrases non-reproduction without claiming a fix, and keeps the writeup at class level with the mitigation.

for a principal

Owns the written policy, the sign-off, and the tradeoff between public pressure on a dependency and arming everyone downstream of it.

## Why the classic trigger has no termination condition Ordinary coordinated disclosure rests on two assumptions: that a fix is a discrete, observable event with a version you can cite, and that the vendor owes the reporter a response as part of a programme. Neither holds for a hosted model. The endpoint can improve or regress with no artefact you can point at; the property may be ruled intended, in which case no fix is ever coming; and your firm may be in no intake relationship at all, having found the behaviour while testing someone else's product. A policy that says "publish when it is fixed" therefore has no stopping rule, and every decision made without one gets made under pressure from whoever wants the write-up out before a conference deadline. So the policy has to be written before the first submission, and it has to answer three questions. ## Trigger: pick conditions you control Legitimate triggers are ones that do not depend on the provider doing anything observable: - a defined interval after a documented filing, whatever the response; - an explicit ruling — fixed, declined, out of scope — quoted with its date; - your own re-measurement showing the effect no longer reproduces at the previously measured rate, stated as an observation and never as an attribution. That last distinction is the one people get wrong. "It no longer reproduces" is something you measured. "They fixed it" is an inference about a system whose internals you cannot see, and it may be false in three different ways at once: a routing change, a temporary safety-layer tightening, or your own harness drifting. ## What publishing responsibly costs Publication is not free, and a policy that assumes it is produces write-ups whose headline number is six months stale. Re-measuring before press time means re-running the suite: 200 attempts with a judge is ~400 calls single-turn, and several times that if the effect is multi-turn — small money, but a day of engineer time to stand the harness back up and re-validate the scoring. The dominant cost is the consent cycle. The findings came out of a client engagement, so the client owns the report, the configuration details, and anything that fingerprints their deployment; getting written sign-off through their counsel is a matter of weeks on their calendar, not yours. Add internal legal review and a named approver. Budget publication as a multi-week track that runs in parallel, and start it at filing time rather than after. ## Where the published number misleads The measured rate is the most quotable and least durable thing in the artefact, and it goes wrong three ways. **Staleness.** Readers quote a figure for years. A hosted endpoint drifts, so an undated rate becomes a claim about a system that no longer exists. Put the measurement date in the same sentence as the number, not in a footnote. **Denominator swap.** You measured over deliberately crafted attempts. Readers hear "this model fails 30% of the time" about ordinary traffic. Write it as "30% of 200 crafted multi-turn attempts under the stated configuration" and let the phrase be clumsy. **Scoring.** The rate is a function of the judge and the harm definition. Another team reproduces the setup with a stricter definition, gets a third of the number, and the gap reads publicly as a refutation of your work. Publishing the criterion — in words a stranger can apply — turns that into a disagreement about definitions rather than about your competence. ## Content limits: the class, not the instance Publish the mechanism, the conditions, the measured rate with its denominator and date, the scoring rule, and the mitigations a reader can apply at a layer they own. Withhold the operative text. The usual argument for full detail — that defenders need the payload to test their own systems — is materially weaker here than in software disclosure, because there is frequently no patched version, so publication does not render the string inert, and the same text works against every deployment of the same dependency rather than against one product. Where the provider has ruled the behaviour intended, the write-up is a claim about a design property, and the wording should say exactly that rather than borrowing the language of a vulnerability. ## Consent, attribution and sign-off The client consents in writing; you strip their identity, their prompts, their data, and any detail that fingerprints their configuration. Never attribute the finding, or the question that produced it, to a company that has not agreed to be named, and never phrase anything so that a reader infers provider endorsement of a conclusion they did not give. Name the approver in the policy, so the decision is a policy applied by an owner rather than a debate held the week of a submission deadline. Two purposes justify publishing at all: helping other builders apply a mitigation they own, and creating collective pressure toward a change no single customer can compel. Both are fully served by a class-level artefact. Neither needs a working recipe.

  • You re-test and the behaviour no longer reproduces. Can you write that the provider fixed it?
    No. You can write that it no longer reproduced at a given date under stated conditions. Attribution to a fix is an inference about a system you cannot see.
  • Why is withholding the exact prompt more defensible here than in classic software disclosure?
    There is often no patched version, so the string does not become inert on publication, and it works against every deployment of the same dependency, not just the one you tested.

saying these in an interview costs you the question

  • Publishing a working prompt because a software-disclosure norm would have published a proof of concept.
  • Claiming the provider fixed it when you only observed it stop reproducing.
  • Publishing engagement detail without the client's written consent.
  • Having no policy and deciding case by case under conference or marketing pressure.
  • Naming the client, or implying provider endorsement of your conclusions.

context