Six weeks after you reported a payload that a hosted content-moderation service failed to flag, the customer says they cannot reproduce it, and the service exposes no version you can pin. How do you resolve the dispute, and how should the finding have been written to survive this?
answer
- no version to pin, use a window
- control set separates drift from noise
- re-run stored payloads verbatim, repeat
- behaviour family beats one string
- hand over a regression corpus
basics
~20 sAssume the service changed behind the endpoint; you cannot pin a version. Re-run the stored payload alongside a control set of cases that flagged during the original run. If the controls still flag and the payload now does too, the behaviour moved. Report that with timestamps instead of defending the old result.
solid answer
~50 sTwo things produce a non-reproducing finding: the service changed server-side, or the response was never stable and the original run caught one draw. Your control set separates them. Controls that still behave as before plus a payload that now flags means drift; controls that also moved means the endpoint, region or configuration differs from the original run. The deeper fix is in how the finding was worded. A finding pinned to one exact string is fragile — a small server-side change retires it, and it also amounts to handing over a reusable payload. A finding phrased as a **behaviour family** ("this class of framing was not covered during the observation window; here are representative cases and their raw responses") survives, because it names the gap rather than one instance of it. Every hosted-filter finding therefore carries an observation window, the raw stored responses, and the repeat count behind it. Those three turn "you were wrong" into "it changed, and here is when".
go deeper
Knows a hosted service can change without notice and that results should be timestamped.
Re-runs the stored payloads with a control set and repeats, and compares raw values rather than pass/fail.
Reads each combination of control and payload outcome correctly, distinguishes drift from nondeterminism from an environment mismatch, and words findings as behaviour families with an observation window.
Makes observation windows, control sets and raw-evidence retention a standing convention, and pushes the customer toward continuous regression testing of controls they cannot version.
**The structural problem.** You are making a claim about a system whose behaviour is controlled by someone else, changed on their schedule, and announced to nobody. A hosted moderation service exposes no version you can pin to a run; even where a model name appears in the request, the vendor may update the model, the policy behind the categories, the calibration, or the routing under that same name. There is no artefact to attach to the finding that fixes what was tested. That is not a reason to avoid the finding — it is a reason to **timestamp it and scope it to an observation window**, which is the only honest unit of claim available. **Why the dispute happens at all.** A non-reproducing finding has exactly four candidate causes, and they are distinguishable if you set the run up to distinguish them: the service changed server-side; the response was never stable and the original run caught one draw; the customer is reproducing against a different environment (endpoint, region, service tier, deployment name, key, parameters, encoding); or the customer is going through their product while you went direct, so the app is transforming the text before it is screened. **The re-verification protocol.** Re-send the stored payloads *exactly* as originally sent — same encoding, same endpoint, same region, same parameters — together with the known-positive **control set**, in one session. Repeat each a few times: a single re-run is precisely as weak an observation as the single run that started the argument. Compare **per-category values, not pass/fail**. A value that moved from just under the cut-off to just over is a different story from one that moved from near-zero to strongly flagged, and the first story may be entirely explained by the customer's cut-off rather than by anything the vendor did. **Reading the outcomes.** | Controls | Disputed payload | Reading | |---|---|---| | Behave as before | Now flags | Server-side change. The finding held in its window — mark it resolved-by-vendor-change, not retracted. | | Behave as before | Still clean | Your setup matches; the customer's differs. Chase environment, tier, or the app transforming text before screening. | | Also changed | Anything | Environment mismatch. Conclude nothing about the payload until the setups reconcile. | | Values oscillate across repeats | Oscillates too | The original observation was noise. Retract it plainly; a clean retraction costs less than a defended coincidence. | **Where the number misleads.** The temptation under dispute is to re-run once, get the old result back, and declare victory. One call that agrees with you is the same evidence as one call that disagrees — and if you did not re-send the controls, you cannot even claim the two sessions were comparable. The second trap is silently re-thresholding your own analysis until the controls line up again: that manufactures agreement and hides the environment mismatch you were supposed to find. **Writing the finding so this rarely happens.** Lead with the **class** of gap, not the instance. Support it with several representative cases rather than one, quote the raw per-category values, and put the observation window and repeat count in the finding itself. A finding pinned to one exact string is fragile twice over: any small server-side change retires it, and it reduces the deliverable to a reusable payload rather than a description of what is not covered — which is the part the customer can actually act on. **The recommendation that outlives the finding.** Hand the customer a small regression corpus of their own, run continuously against the filter, alerting on movement in the per-category values. A control nobody can version can regress in either direction between engagements — quietly getting worse, or quietly getting stricter and eating legitimate traffic — and without continuous measurement neither is visible until a user complains. That recommendation is frequently worth more than the finding that provoked it, and it costs the customer a scheduled job and a small metered budget rather than another engagement.
- Your re-run shows the control set also behaving differently. What does that tell you?That the environments differ — endpoint, region, tier, key or parameters — so nothing about the disputed payload can be concluded until the setup matches the original run.
- What do you recommend the customer do so this class of dispute stops recurring?Own a small regression corpus that runs against the filter continuously and alerts on movement, since an unversioned hosted control can regress silently between engagements.
Testing an unversioned hosted filter is like reporting the depth of a river: correct on the day you measured, and worth nothing unless you wrote down the date. The control set is the fixed gauge on the bank that tells you whether the river moved or your ruler did.
saying these in an interview costs you the question
- Defending the original result without re-running anything.
- Re-running once and treating that single call as proof either way.
- Findings with no timestamp, no observation window and no stored raw response.
- A finding that consists of one exact payload string and nothing else.
- Blaming the vendor for a non-repro that is actually a different region, tier or endpoint.