skip to content

The same jailbreak framing works against other products on the same hosted model - what does that tell you about the defect you filed?

level: seniorimportance: nice to knowfreq 33%

answer

  1. hold everything constant but the model
  2. the shared variable is the weights
  3. shipping a change makes you no safer than rivals
  4. location is not exposure
  5. asymmetric results are the actionable ones

basics

~20 s

It locates the defect. The one variable held constant across those products is the model, so the behaviour is model-layer rather than something your prompt or surface created. Your exposure stays yours: the transcript still carries your product's name.

solid answer

~50 s

Reproducing across unrelated products that share a supplier's model is a locating measurement, not a curiosity. Everything else differs - system prompts, interfaces, user populations, how each app renders what comes back - and the framing still lands, so the behaviour belongs to the shared weights rather than to anything your team wrote. Three things follow. Changes on your side can only move the observed rate, because the artefact producing the behaviour is not in your repository. Your finding is not a competitive differentiator: nobody on that model is safer than you after you ship a prompt change. And the exposure is still yours, because the person who saw the output was your customer and the transcript carries your product's name. Watch the direction of the claim: shared behaviour does not mean equal exposure, since what each product does with the returned text differs, and that part is per-product.

go deeper

for a junior

Know that many products call the same hosted model, so behaviour seen in one can appear in all of them, and that this is about where the behaviour comes from rather than who is affected.

for a middle

Explain why testing a second deployment is a controlled experiment: everything but the model differs, so a landing framing implicates the shared weights.

for a senior

Demonstrate the two-attribution write-up - generating behaviour to the shared model, consequence to your own product - and show you treat an asymmetric result as the more actionable one.

for a principal

Own the framing for stakeholders who hear 'everyone has it' as 'nobody needs to act', and keep exposure accounted for on your side even when the fix is not available anywhere in your estate.

**Testing a construction against a second deployment is the cheapest experiment in this whole area, and it answers the question triage actually needs answered: which half of the estate is this in?** ### The measurement You have a filed transcript from a study-companion chat app: a user's own typed framing, and content the model normally declines. Try the same class of framing against another product built on the same hosted model - a different team, a different system prompt, a different interface, different customers. If it lands there too, then across the two runs almost every variable differed except one: the model. Attributing the behaviour to the shared component is ordinary controlled reasoning, and it is far more convincing than any argument about the wording itself. ### What follows, and what does not **It follows that app-side changes cannot remove it.** The artefact producing the behaviour is not in your repository, so anything you ship can only move how often it happens for the wordings you measured. That is worth doing and worth recording, but it is not removal, and the ticket should not say otherwise. **It follows that this is not a differentiator.** A team sometimes treats a jailbreak finding as a competitive gap to be closed. When the behaviour is shared by every deployment of the model, shipping a change does not make you safer than the product next door in any way a customer can perceive; it changes a rate on one surface. **It follows that the report has a second useful recipient.** The party who can change the behaviour is the one who trains and serves the model. Sending it there is a different act from fixing it, and it does not transfer your exposure. **It does not follow that the finding is not yours.** This is the inversion to guard against. Location and exposure are different things. The person who saw the content used your app, the transcript carries your product's name, and if that output reaches a minor using a study companion, nobody is comforted that the same thing happens elsewhere. A defect you cannot patch is still a defect you are accountable for reporting honestly. **It does not follow that every product is equally exposed.** Shared trained behaviour is the input; what each product does with the returned text is not shared. Two apps on the same model can differ enormously in what the output reaches, who sees it, and what it is presented as. Consequence is per-product even when the generation is not. ### The asymmetric case is just as informative Suppose it lands on your product and not on the other. Now the shared component is ruled out as a sufficient explanation, and something in the differing half participates - the way your system prompt frames the assistant's role, the surface the user is on, how turns are assembled. That half has a repository, a build and a date, so a finding that reproduces asymmetrically is the more actionable of the two, even though it looks like the weaker result. Many engineers read it the other way round and chase the more spectacular cross-product result while ignoring the one they could actually work on. ### How to write it up The useful record separates the location from the ownership. State the construction as a class, not a string. State which deployments it was observed against and how many attempts each took. State that the generating behaviour is attributed to the shared model, and that the consequence - what the output reached, who saw it, what the product's presentation implied about it - is attributed to your own product. Two attributions, one finding. Collapsing them into a single owner is what produces both of the bad outcomes: 'the supplier's problem, nothing for us to do', and 'our bug, we will fix it next release'.

  • Does a shared, model-layer behaviour mean the finding is not your team's problem?
    No. Location and exposure are separate. The behaviour lives in weights you do not control, but the output reached your customer, on your surface, under your product's name. Your team still owns what the product did with that text, who could see it and how it was presented - and owns saying plainly, in the record, that the generating behaviour is unchanged by anything you ship.
  • What would you conclude if it lands on your product but not on another built on the same model?
    That the shared model is not a sufficient explanation and something in the differing half participates - how your system prompt frames the assistant, which surface the user is on, how turns are assembled. That half sits in a repository, so this asymmetric result is the more actionable one: it is the version of the finding that can honestly carry an owner and a date.

saying these in an interview costs you the question

  • Concludes a shared defect is nobody's finding
  • Treats an app-side change as making the product safer than rivals
  • Assumes equal exposure across products on one model
  • Ignores an asymmetric result as a weaker finding
  • Says the supplier owns it, so there is nothing to record

context