skip to content

You are filing a model-behaviour report with the provider of a hosted chat endpoint your client builds on. What evidence does that submission need that a report of a deterministic software bug would not, and why?

level: middleimportance: should knowfreq 44%

answer

  1. rate, not a transcript
  2. attempts over successes
  3. model id as returned
  4. temperature, seed, system prompt
  5. who scored a hit
  6. timestamp: the endpoint moves

basics

~20 s

The behaviour is probabilistic, so one transcript is not a reproduction. Give how many attempts you made and how many succeeded, the exact endpoint and model identifier, the decoding settings and any system prompt, and the date and time window. Without a rate and those conditions, the reader cannot separate your result from noise.

solid answer

~60 s

A deterministic bug report says "do this, get that". A hosted model gives you a distribution, so the equivalent claim is a **rate under stated conditions**. Send: attempts and successes (not just a winning transcript), the endpoint and model identifier exactly as the API reports it back, sampling parameters such as temperature and any seed, the full system prompt and message sequence, and when you ran it. Two things make this harder than it sounds. First, your own scoring is a judgement call — if an automated judge decided which responses counted as hits, say what the judge was and what it treated as a hit, because the provider may score it differently and get a lower rate. Second, hosted endpoints move underneath you: routing, safety layers and weights can change without a version bump you can cite, so a timestamp is part of the evidence, not decoration. Also state what you did *not* vary. A rate measured at one temperature with one system prompt says nothing about the default configuration a provider will retest with.

code

yaml · 11 lines
yaml
observed_effect: <one sentence, what the model did>
attempts: 50
successes: 11
success_criterion: <who or what decided a response counted, and by what rule>
endpoint: <host / route, as configured>
model_identifier: <exactly as returned in the API response>
sampling: { temperature: ..., top_p: ..., max_tokens: ..., seed: ... }
conversation_shape: <single-turn | multi-turn, number of turns>
system_prompt_present: <yes/no, summarised>
run_window: <start and end timestamps, timezone>
not_varied: <parameters held fixed, so the claim's limits are explicit>

go deeper

for a junior

Knows to attach the prompt, the response and the model identifier rather than a screenshot.

for a middle

Reports successes over attempts with the sampling settings and conversation shape, and explains why one transcript is not a reproduction.

for a senior

Discloses the scoring rule, keeps raw run records so the rate can be recomputed under the provider's definition, and re-runs before escalating.

for a principal

Standardises the evidence schema across the team so submissions and customer reports quote comparable rates measured the same way.

## Repro procedure versus repro statistic A deterministic software bug report carries a *procedure*: run this input, observe this output, every time. A hosted chat endpoint samples from a distribution, so the equivalent claim is a **rate under stated conditions**, and everything the submission needs follows from making that rate checkable by someone who was not there. The conditions block, and why each entry exists: | field | why the triage team needs it | |---|---| | attempts and successes | a single transcript is unfalsifiable; a fraction is a claim someone can re-measure and argue with | | model identifier as returned by the API | the response records what actually answered; a service may route among variants, so the request only records your intent | | sampling settings (temperature, top-p, max tokens, seed) | a rate at temperature 1.2 and a rate at the service default are numbers about different systems | | full prompt state (system prompt, tool definitions, prior turns) | multi-turn effects vanish when a reviewer replays only the last message | | the success criterion | if a classifier or model judge labelled the outputs, the rule it applied *is* the definition of the rate | | run window and timezone | the endpoint can move under you, so the number has an expiry date | | what you did not vary | it states the claim's limits before someone else overstates them for you | ## What a defensible submission costs Fifty single-turn attempts is fifty calls — trivially cheap. The real bill starts when the effect needs a build-up: a six-turn sequence at 50 attempts is ~300 calls, an automated judge adds one call per attempt, and you generally need a second arm measured at the provider's *default* configuration so the comparison exists before they ask for it. That is realistically 800-1,500 calls for one finding: single-digit to low-tens of US dollars on most hosted endpoints, an hour or so of wall clock against a per-minute rate limit with retries, and — the dominant cost — most of an engineer-day building the harness and hand-labelling a sample of transcripts to check that the judge agrees with a human. ## Where the number misleads **Sample size.** A rate from 20 attempts is nearly uninformative. Five successes in 20 is a 25% point estimate whose 95% interval runs roughly 9% to 49%; 25 in 100 is the same 25% with an interval of roughly 17% to 35%. If the finding's weight rests on "this happens often", 20 attempts cannot carry it, and a provider who re-runs 20 of their own and sees one success has not contradicted you. **The judge.** This is the most common way a submitted rate is simply wrong. Suppose the true rate is 10% and the automated judge agrees with human labels about 85% of the time. Of 200 attempts, ~20 are real successes and ~180 are not; a judge that mislabels roughly 15% of the negatives contributes on the order of 27 false positives, so the reported count is dominated by judge error and the headline rate is inflated well past 20%. Nothing about the transcript is fake — the *scoring* is. Hand-label a random sample, publish the agreement number alongside the rate, and keep the raw run records so the rate can be recomputed under whatever definition the provider comes back with. **Denominator.** Your rate is over deliberately crafted attempts. A reader hears it as a property of ordinary traffic. Write the denominator into the same sentence as the number, every time. **Drift.** The system you measured may not exist next week, in either direction, which is why the timestamp is evidence rather than metadata. ## The "cannot reproduce" reply, and what to check Most non-reproductions are conditions mismatches rather than disagreements about reality: the provider retested single-turn, at the default sampling configuration, with no system prompt, at a smaller attempt count, and scored a hit under a stricter harm definition. If your submission already stated all five, the exchange becomes a negotiation over the rate and the criterion instead of an argument about whether anything happened. Before sending, check: are attempts and successes both present; is the identifier the one the API returned; is the scoring rule disclosed in words a stranger could apply; did you sanity-check the judge against human labels; is the attempt count large enough for the claim you are making; and have you re-run recently enough that the number is still about the live system? Withhold the copy-paste operative text — the class, the conditions, the rate and the criterion are what make the report actionable, and the payload itself is not needed for triage.

  • The provider replies that they cannot reproduce it. What is your first hypothesis?
    A conditions mismatch: they retested single-turn, at the default sampling configuration, without your system prompt, or scored a hit more strictly. Compare conditions before re-arguing the effect.
  • Why record a timestamp for a model-behaviour finding?
    Hosted endpoints change without a version you can cite. The timestamp is what lets anyone say whether a later non-reproduction means you were wrong or means the system moved.

An automated judge that agrees with humans about 85% of the time, scoring a behaviour that only happens in one attempt out of ten, is like a smoke alarm that also goes off for toast: most entries in the log are toast, so the total says more about the alarm than about the fire.

saying these in an interview costs you the question

  • Submitting one winning transcript and calling it reproducible.
  • Not recording the model identifier the API actually returned.
  • Hiding the scoring rule that decided which responses counted as successes.
  • Reporting a rate measured at an unusual temperature as if it were the default behaviour.
  • Discarding the raw run records, so the rate cannot be recomputed under a different definition.

context