skip to content

A chat model declines one variant of an attacker's probe and answers another: what does either single observation prove?

level: middleimportance: should knowfreq 55%

answer

  1. one draw is not a rate
  2. replies are sampled, not computed
  3. repeat before you believe it
  4. spend repeats where outcomes disagree

basics

~20 s

Very little alone. Replies are sampled, so one decline shows only that this draw refused and one answer only that this draw did not. The boundary is a rate, and estimating a rate costs repeated probes.

solid answer

~50 s

Each reply is one draw from a distribution, so a single observation is a sample, not a measurement. A decline establishes that the reply produced on that occasion was a refusal; it does not establish a rate, and it certainly does not establish that the neighbourhood is covered. The same is true in reverse: one answered probe is not evidence that the topic is open. To get anything usable, the mapper has to treat each probe as repeated trials and estimate a decline rate, which is where the cost lands — on a consumer product with a per-account message quota, every repeat is a message not spent on breadth. The useful discipline is to spend repeats where outcomes actually vary, since a probe that has declined every time so far returns almost no new information for the next message.

go deeper

for a junior

Know that replies are sampled, so the same message can be declined once and answered once, and that one observation is therefore not evidence about the model in general.

for a middle

Be ready to explain the mechanics: outcomes are draws from a distribution, so the honest unit is a decline rate with a sample count, and repeats are what buy it.

for a senior

Show the operating judgment — a control probe to detect the system moving mid-pass, repeats concentrated where outcomes disagree, and a refusal to promote a single lucky draw to a result.

for a principal

Own the consequence for what gets claimed: any number leaving your team should carry its sample count and its window, because an unqualified rate invites decisions the evidence does not support.

## Why the same input gives different replies Generation samples. Unless a deployment pins decoding to the single highest-probability token — and consumer chat products generally do not — the reply is drawn from a distribution over continuations. Refusal lives inside that distribution rather than beside it, so an input near the boundary can produce a decline on one draw and an answer on the next without anything having changed. And sampling is not the only source of movement. The system behind the chat box is not guaranteed to be the same artefact from one probe to the next: a deployment can change, and the surrounding instructions the product wraps around the conversation can change too. From outside, none of that is announced. All the prober sees is that an input which behaved one way an hour ago behaves differently now. ## What a single observation licenses Precision about direction matters more here than anywhere else in a mapping pass. | Observation | What it establishes | What it does not | | --- | --- | --- | | One decline | this draw produced a refusal | that the neighbourhood is reliably covered | | One answer | this draw did not refuse | that the neighbourhood is uncovered | | One hedge | this draw sat between the two | any rate at all | Every row says the same thing: the unit of evidence is a rate, and a rate needs repeats. An attacker who records binary outcomes from single probes is building a map out of coin flips, and will read noise as structure — usually as "this topic is open" on the strength of one lucky draw, which then fails to reproduce for anybody else. ## Where the information actually is A probe that has declined on every one of ten attempts is deep inside covered territory; the eleventh attempt tells you almost nothing. A probe that has answered every time is deep outside it; same conclusion. The probes worth repeating are the ones whose outcomes disagree with each other, because that is where the decline rate sits away from the extremes and where an extra sample moves the estimate most. That is the whole allocation principle: sweep wide with single probes to find disagreement, then concentrate repeats on the disagreeing band. ## Why this is expensive rather than merely annoying On a consumer chat product the vantage point is an account and the account is metered. Every repeat is one message from a finite quota, so resolution and breadth trade directly against each other: ten repeats each across twenty neighbourhoods costs the same as one probe each across two hundred. There is a second cost too — an account that burns its quota on obviously repetitive probing is an account that may not survive to finish the map, and the map dies with the vantage point. ## The control probe Because the deployment can move underneath a mapping pass, a mapper carries at least one probe whose outcome is already well established and re-issues it periodically. If the control starts behaving differently, the change is in the system, not in the new probes, and everything measured since the last good control is suspect. Without a control, a mid-pass change looks exactly like a discovery. ## The misreadings to avoid The first is treating one decline as proof of coverage — the same crisp-boundary assumption that underlies the banned-list picture, just applied from the other side. The second is treating one answer as a finding: a construction that worked once against one deployment on one draw is not yet a reliable result, and reporting it as one is how a red-team report acquires items that nobody else can reproduce. The third is attributing variance to the wrong cause: before concluding that a phrasing difference moved the outcome, the mapper needs enough samples on both phrasings to show that the difference is bigger than the draw-to-draw noise. What survives all of this is modest and useful: a decline rate per probe, with a sample count attached, dated to a window. That is the honest unit a coverage map is built from.

  • Besides sampling, what else can make the same probe behave differently an hour later?
    The system behind the chat box need not be constant. The deployment can change, and the instructions the product wraps around the conversation can change, neither of which is announced to a user. From outside they are indistinguishable from noise, which is why a mapper re-issues a known control probe periodically — if the control moves, the change is in the system, not in the new probes.
  • Where would you not spend another repeat?
    On a probe that has produced the same outcome every time so far. It already sits well away from the boundary, so the next sample barely moves the estimate. The repeats belong on probes whose outcomes disagree with each other, since those sit near the middle of the rate and carry the most information per message spent.

Asking once whether a shop is open tells you about that moment, not its hours. A schedule needs repeat visits, and the visits worth repeating are the ones that came back inconsistent.

saying these in an interview costs you the question

  • Treats a single decline as a measurement of the boundary
  • Calls a one-off successful probe a reproducible finding
  • Assumes the deployment is identical between probes
  • Spends repeats uniformly instead of on disagreeing probes
  • Attributes draw-to-draw noise to a phrasing difference

context