skip to content

The Half You Control

A finding against the model itself carries no deterministic reproduction and no version to fix it in, and the same framing works wherever that model runs. Interviewers listen for a promised fix date.

on this pageshow

explore

questions

4

A user's typed framing got your chat app's hosted model to emit content it normally refuses - which half of that defect can your release fix?

level: juniorimportance: must knowfreq 60%

answer

  1. ask which artefact would change
  2. your repository, or a supplier's
  3. the prompt is yours, the weights are not
  4. only one half has a version to bump

basics

~20 s

Only the half you ship. Your repository holds the system prompt, the surfaces and what the app does with returned text. The refusal that framing got past is trained behaviour in a supplier's hosted model, with no release of yours.

solid answer

~50 s

Split the finding by asking which artefact would have to change for it to stop. In a tool-less study-companion chat product, the things your team ships are the system prompt, the chat surface, session handling and whatever the app does with the text that comes back - all of that has a repository, a build and a date. The propensity to decline a request lives in weights you did not train, served from an endpoint you do not operate. The OWASP Top 10 for LLM Applications spans both halves: improper output handling, excessive agency and system-prompt leakage describe your own code around the model, while a framing that gets past a trained refusal describes the model's behaviour. Changes on your side can move how often it happens; none of them removes the behaviour, and that is what decides whether a remediation date is honest.

go deeper

for a junior

Be ready to say which parts of an LLM product your team actually ships: the system prompt, the interface, and what the app does with the returned text. The model's willingness to refuse is not one of them.

for a middle

Explain why a prompt reword cannot remove the behaviour - the prompt is text the model reads, not a rule the runtime enforces - and which entries on the LLM list describe your own code rather than the model's.

for a senior

Show that you split a finding by artefact before you triage it, and that you describe an app-side change by what it measurably shifts rather than by what it closes.

for a principal

Own the consequence for the process: half these findings have no release of yours to land in, and treating them as if they do costs the credibility of every date your organisation gives.

**Every finding has to answer one question before it can be scheduled: which artefact would have to change for this to stop happening?** On an LLM product that question splits the estate in two, and only one half has a version number. ### The shape of the product decides the answer Take a study-companion chat app on a phone. Its team writes a system prompt, the user types turns, and a hosted frontier model answers over an API behind a version label the vendor rotates. There is no retrieval, no tool, no memory outliving the session. The only text reaching the model that the team did not write is the paying customer's own message, and the only thing that can leave the system is text rendered back to that same customer. That shape removes one of the two things people usually mean by prompt injection. Injection aims at the *application's* own instructions, and it needs a path by which text somebody else wrote reaches the model as part of the application's data - a retrieved chunk, a fetched page, an inbound mail, a tool's return value. This product has no such path. What is left is the other target: the *model's* trained refusal, addressed by the user in their own turn. The published list files jailbreaking under its prompt-injection entry, so one entry spans both targets - and the ownership line runs straight through the middle of it. ### Split the estate | What it is | Who ships it | A version you can bump? | | --- | --- | --- | | The system prompt text | your team | yes | | The chat surface, session handling, client | your team | yes | | What the product does with the returned text | your team | yes | | Which model and which label you call | your team | yes - you can change supplier, not their behaviour | | The propensity to decline a request | the supplier | no | | How that propensity is sampled at generation time | the supplier | no | Everything above the line has a repository, a build and a release you can put a date on. Everything below it is trained behaviour in weights you did not produce, served from an endpoint you do not operate. Several entries on the LLM list sit squarely in the first group - how the application handles what the model returned, how much the application is allowed to do on the model's say-so, what the team put in the system prompt that later leaked. Those findings have an owner, a repository and a date. A framing getting past a trained refusal does not, at least not on your side of the line. ### What an app-side change actually buys Rewording the system prompt is a change in the half you control, and it can move the observed rate for the wordings that were tested. It does not alter the trained propensity, because the propensity is not stored in the prompt - the prompt is text the model reads, not a rule the runtime enforces. The same goes for narrowing which surfaces exist or which users reach them. The honest description of any such change is what it measurably shifts, never what it removes. Describing it as removal is how a team ends up telling a reviewer that a defect is closed while the same transcript is still reproducible. ### Why this matters more than it sounds The usual wrong answer is 'we will fix it in the next release'. Your next release ships the prompt, the surface and the client. It does not ship the weights. A date on somebody else's release train is a promise nobody in your organisation can keep, and the cost is not only this ticket: the next date your team gives is believed a little less. ### Getting the direction of the claims right One success proves the framing landed once, on one deployment, at one moment. A refusal on the next attempt proves that attempt was declined, not that the finding evaporated. And 'we cannot patch it' is not the same as 'it is not our problem': the transcript carries your product's name, the person who saw the output used your app, and the exposure is yours even when the behaviour is not.

  • The team proposes rewording the system prompt to stop it. What does that buy them?
    A change in the half they control, with a build and a date. It can move the observed rate for the wordings they tested, because it changes the text the model reads. It does not touch the trained propensity, so the class of framing survives it. The honest ticket entry is what the change measurably shifted, not that the behaviour was removed.
  • A reviewer files it as 'a bug in the model'. Is that the right framing?
    It is imprecise in a way that stalls the ticket. There is no component to point a patch at: the behaviour is a statistical property of a trained system a supplier serves. File what was observed against your product, attribute the model half to the supplier's endpoint, and keep the two halves separate so the half with a repository can be worked and dated.

It is the difference between a bug in your code and a bug in the platform you rent. You can change how your app calls it and what it does with the answer; you cannot ship a version of somebody else's runtime.

saying these in an interview costs you the question

  • Says it will be fixed in the next release
  • Believes the refusal is configured in the system prompt
  • Assumes every entry on the LLM list maps to your own code
  • Treats a trained propensity as a runtime setting
  • Concludes the finding is invalid because you cannot patch it

context

open as a page

A filed jailbreak against your hosted chat model reproduces once in five tries - what does that prove?

level: middleimportance: should knowfreq 52%

basics

~20 s

It proves the framing landed once, against one deployment, at one moment. Refusal is a trained propensity sampled at generation time, so a failed retry means that attempt was declined - not that the finding is wrong or the behaviour gone.

open as a page

Triage wants an owner and a fix date for a jailbreak in a supplier's hosted model - what do you commit to?

level: principalimportance: should knowfreq 32%

basics

~20 s

Commit only on the half with a repository: name what your team will ship and when, describe it by what it measurably shifts, and record the model-side behaviour as unresolved rather than scheduled. A supplier's release train takes no date from you.

open as a page

The same jailbreak framing works against other products on the same hosted model - what does that tell you about the defect you filed?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

It locates the defect. The one variable held constant across those products is the model, so the behaviour is model-layer rather than something your prompt or surface created. Your exposure stays yours: the transcript still carries your product's name.

open as a page