Mapping a chat model's refusal coverage from one free-tier account: how do you spend a finite message quota?
answer
- coarse sweep before deep repeats
- the transcript is part of the input
- move one dimension per probe
- carry a probe whose outcome you know
basics
~20 sCoarse first, then deep. Sweep neighbouring topics with one probe each to find where outcomes vary, then concentrate repeats on that band. Isolate probes in fresh conversations, move one dimension at a time, and carry a control probe throughout.
solid answer
~50 sTreat it as a measurement campaign, not a hunt. First pass: one probe per neighbourhood across a wide sweep, recording the reply's shape rather than a yes/no — flat decline, hedge, partial answer, redirect, full answer. Second pass: spend repeats only where the first pass disagreed with itself, because that is where the decline rate sits away from the extremes. Three disciplines make the results mean anything. Issue each probe in a fresh conversation, since earlier turns are part of the input and a probe sent after a decline is no longer the same probe. Move one dimension between probes, so a difference in outcome can be attributed to something. And re-issue a settled control probe periodically, so a deployment change is detected rather than recorded as a discovery. Budget the account itself too: the vantage point is one account, and the map dies with it.
go deeper
Know that a probe's context matters: the same message sent into a fresh conversation and into one that already contains a decline are two different inputs.
Be ready to explain the two-pass shape — breadth first, then repeats only where outcomes disagreed — and why changing one dimension per probe is what makes a difference attributable.
Demonstrate campaign hygiene end to end: controls against a moving deployment, ordinal reply shapes instead of a boolean, and an allocation decided before the quota starts draining.
Own the reporting posture — a map is one account, one deployment, one window, and the claim you let leave the team must not outgrow that, because stripped qualifiers are how a snapshot becomes a false assurance.
## What the campaign is trying to produce The deliverable is a ranked picture of neighbourhoods — which related subjects the refusal propensity covers thickly, which thinly, and which not at all — dated and scoped to one deployment seen through one account. It is not a set of working constructions, and it contains no content the product would not otherwise emit. Its value is aim: later effort spread evenly over an uneven boundary mostly lands in covered territory and returns nothing. ## Two passes, not one A single pass at uniform depth is the common mistake. It either samples too shallowly to distinguish structure from draw-to-draw noise, or samples deeply everywhere and covers a fraction of the ground. **Pass one, breadth.** One probe per neighbourhood, spread across the space of related subjects, recording what came back. **Pass two, depth.** Repeats concentrated on the neighbourhoods where pass one produced inconsistent or intermediate replies, because those sit near the middle of the decline rate and are where an extra sample is worth most. A neighbourhood that declined flatly and one that answered fully both need very little more attention. ## Record shape, not a boolean Replies are not binary. A flat decline, a decline with an offer of something adjacent, a heavily caveated partial answer, a redirect to a different subject, and a full answer form an ordered set, and the ordering is the gradient of the boundary. Collapsing them to declined/answered discards exactly the information that distinguishes a thin region from an uncovered one. Record the shape; rank later. ## Isolation: the transcript is part of the input A probe issued in a conversation that already contains a decline is not the same probe as one issued into an empty conversation — the earlier turns are in the model's input and change what is being measured. A mapping pass that reuses one long thread produces a map of that thread, not of the product. So every probe gets a fresh conversation, which costs whatever the product charges for starting one, and that cost belongs in the budget from the start rather than as a surprise halfway through. ## One dimension at a time If two things differ between consecutive probes, neither can be credited with a difference in outcome. Hold the request fixed and move only the neighbourhood; or hold the neighbourhood fixed and move only one property of the request. This is unglamorous and it is what separates a map from a pile of anecdotes — and given that outcomes are sampled anyway, an unattributable difference is indistinguishable from noise. ## Controls, because the ground moves The system behind the chat box is not guaranteed constant for the length of an engagement. Carrying one or two probes whose outcome is already well established, and re-issuing them at intervals, is the only outside way to notice. If a control moves, everything measured since the previous good control is suspect and the map needs a date boundary drawn through it. Without controls, a deployment change reads as a finding. ## Budgeting the account, not just the messages On a consumer product the account *is* the vantage point. Two limits bite. The metered quota sets the total number of observations available, which is why breadth and resolution trade against each other so directly. And the account's own standing is finite: obviously repetitive probing is visible from the inside, and losing the account ends the campaign with whatever was recorded so far. Both argue for deciding the allocation up front — how many neighbourhoods, how many repeats on the disagreeing band, how many messages held in reserve for controls and for re-checking anything surprising. ## What the finished map may and may not be used for It supports statements of the form "in this window, through this account, on this deployment, replies in neighbourhood X declined roughly this often, and neighbourhood Y showed inconsistent coverage across N samples". It does not support "the product refuses X" or "the product allows Y". Keeping the claim at the size of the evidence is what makes the map usable later rather than an artefact somebody quotes back with the qualifiers stripped off.
- Why not run the whole pass inside one long conversation to save on setup?Because earlier turns are part of the model's input. A probe issued after a decline, a hedge, or a long unrelated exchange is a different input from the same probe in an empty conversation, so the outcomes are not comparable. A thread-based pass maps that thread. Fresh conversations cost more messages, and that cost belongs in the budget from the start.
- You have half the quota left and the first pass came back almost entirely consistent. What now?Widen rather than deepen. Consistency everywhere means the sweep did not cross the boundary, so more repeats on settled probes buy nothing. Move the sweep into neighbourhoods further from the ones already covered, or vary a different property of the request, and only start spending repeats once outcomes begin to disagree.
- What do you do with the map if a control probe moves halfway through?Draw a date boundary. Everything measured after the last good control is measured against a system that may differ, so it cannot be merged with the earlier half as if it were one dataset. Re-run a sample of the earlier neighbourhoods to see whether the shift is broad or local, and report the two halves as two dated snapshots.
saying these in an interview costs you the question
- Runs every probe in one long conversation
- Changes topic and phrasing together between probes
- Records outcomes as a bare declined/answered boolean
- Spends the whole quota at uniform depth
- Never re-checks a settled probe during a long pass