A quarter ago your report recorded that a gradient-optimised suffix produced the disallowed behaviour on roughly two of five attempts against a hosted chat endpoint. Re-running the same frozen string today, it almost never works. What are the plausible causes, and what should the original entry have recorded so this is diagnosable at all?
answer
- transferred rate is perishable, stamp the date
- new build, new guard, new system prompt, or noise
- raw counts not a bare percentage
- store failure text verbatim to attribute later
- keep the objective, not just the string
basics
~20 sThe endpoint changed under you: a new model build, an altered system prompt, or a filter added in front or behind it. Your string did not decay. The entry should have stamped the date, the endpoint and any build identifier returned, decoding settings, the number of attempts, and how a hit was judged. Without those, nothing is attributable.
solid answer
~50 sA transferred success rate is **perishable**. It describes a deployed system on one day, and a hosted endpoint is a moving pipeline: model builds roll forward, system prompts get edited, classifiers get added or retuned. Causes worth separating: the model behind the endpoint changed; a guard was added or tightened, so the request never reached the model; the hidden system prompt changed the framing your string relied on; or the original number was never solid — two hits in five attempts is a tiny sample. What the entry needed: measurement date, endpoint identifier and any build string the API returns, decoding settings, raw attempts and hits, the exact judging criterion, and stored samples of the *failure* responses. That last item is what separates "model refused" from "something in front of the model rejected it" a quarter later. Going forward: keep the objective and surrogate so you can re-derive a string, rather than treating one string as the finding.
go deeper
Should recognise that the endpoint can change and that the string itself did not change.
Should list the distinct causes — model build, added guard, altered system prompt, small-sample noise — rather than assuming a fix was shipped for this attack.
Should name the record fields that make it attributable after the fact, and use stored failure text to separate a guard rejection from a model refusal.
Should set the program rule: transferred results carry a measurement date and a re-test cadence, and findings are written at the behaviour-class level so they outlive any one string.
**Start by ruling out the thing that did not happen.** The string did not decay. It is inert text; it has no state, no expiry, no wear. Everything that changed is on the other side of the API boundary, and a hosted endpoint is a product under continuous change: model builds roll forward without renaming the endpoint, system prompts get edited, classifiers get added and retuned, routing sends traffic to different capacity. Four quite different changes produce the identical symptom you are looking at. 1. *A new model build.* Same endpoint name, different weights. The refusal geometry the string was fitted to no longer exists. 2. *A guard appeared or was tightened.* An input classifier is very good at exactly the thing an optimised string looks like — anomalous token sequences — and rejecting it costs the vendor a fraction of a generation. An output classifier suppresses the completion after the fact. Neither is the model changing its mind. 3. *The hidden system prompt changed.* Your string was wrapped in text you never saw; edit that wrapper and the framing the string exploited moves out from under it. 4. *The original measurement was noise.* This one is the most uncomfortable and the most common. **Where the number misleads: two of five is not forty per cent.** "Roughly two of five attempts" is a sample of five. The uncertainty around that is enormous — a handful of Bernoulli trials is compatible with a true rate well under ten per cent and with one well over seventy. Quoting it as "40%" launders a coin-flip's worth of evidence into a figure that looks like a measurement, and it means today's "almost never works" may not be a change at all: it may be the second sample of a rate that was always marginal. If nobody wrote the attempt count down, you cannot even test that hypothesis, and you will spend engineer days chasing a vendor change that never happened. That wasted diagnosis is the real cost of an under-instrumented entry — the queries were pennies; the quarter-later archaeology is not. **What the original entry needed.** Treat every transferred result as a dated experiment, and record: the measurement date; the endpoint identifier plus any build or version string the response metadata carried; the decoding settings you sent and which you could not control; attempts and hits as raw counts, never a bare percentage; the judging criterion — what text counted as the behaviour, and whether a human, a rule or a model decided; and verbatim samples of both successes and failures. That last field is the one people skip and the one that does all the work later. **Diagnosing with that in hand.** Compare today's failure text against the stored failures: | what you observe now | most likely layer | cheap confirming check | |---|---|---| | identical canned sentence, fast, nothing streamed | classifier in front of the model | send a benign message sharing the string's odd surface form | | fluent in-voice refusal, wording varies per attempt | the model itself | vary decoding; refusals should still vary | | output starts then stops mid-completion | output-side check | compare truncation point across attempts | | build or version identifier differs from the recorded one | new model build | re-run a stable non-adversarial probe and compare style | If none of these is distinguishable because nothing was stored, the finding is that the report was under-instrumented — write that down and fix the template, because it will recur. **What actually survives the decay.** Not the string, and therefore not the exhibit. What survives is the *objective and the surrogate* that produced it: keep those, along with the search settings and the held-out results, and you can re-derive a candidate against the new build for the price of GPU hours instead of restarting the whole method. Write the finding at the level of the behaviour class, attach the string as a dated exhibit with raw counts, and state an explicit re-test cadence — otherwise the number will be re-quoted as current by someone reading the PDF next year, which is how a stale rate becomes an organisational belief.
- Today's failures are an identical canned policy sentence returned quickly with no partial output; a quarter ago the failures were varied, in-voice refusals. What changed?Most likely a classifier in front of the model now rejects the request before generation. Uniform wording, no streamed content and low latency are the signature of a separate guard rather than the model declining.
- How do you write the finding so it does not go stale the same way next quarter?State it as a behaviour class demonstrated on a given date against a given endpoint build, attach the string as a dated exhibit with raw counts, and give an explicit re-test cadence plus the objective and surrogate needed to re-derive a candidate.
saying these in an interview costs you the question
- Concluding the vendor 'fixed the vulnerability' with no evidence about which layer changed.
- Re-quoting the old rate as current because it is in a published report.
- Reporting a percentage without the attempt count behind it.
- Not keeping failure responses, leaving no way to distinguish a guard rejection from a model refusal a quarter later.
- Treating the individual string as the finding rather than the behaviour class it demonstrated.