skip to content

A listing written at a generative screen passed once in five re-runs: is that a finding?

level: seniorimportance: nice to knowfreq 29%

answer

  1. one pass, one stage, one version
  2. which stage decided is the first gap
  3. five trials is a very wide interval
  4. rate matters against retry cost
  5. does it survive rewording, or only that wording

basics

~20 s

Yes, if you can say which stage passed it and that the rate beats the stage's own noise. One pass in five proves the construction worked against one deployment at one version — and a retryable path makes a fifth material.

solid answer

~50 s

Start by establishing what the single pass is evidence of: a specific stage, at a specific version, emitting a passing verdict once. Before it is a finding you need three things. Which stage decided — with a label-only classifier in front and a generative screen behind, the record may carry a final outcome and no per-stage decision, so the report has nothing to attach. Whether both stages saw the same bytes the run submitted, since normalisation and truncation sit in between. And whether one in five is above the stage's baseline variation for comparable items, which means a control arm, not a single number. Then judge it on the attacker's economics rather than yours: a delivery path that can be retried at near-zero cost turns a twenty-percent pass rate into reliable passage, and a wording-specific effect that disappears when the wording changes is a much weaker claim than one that survives rephrasing.

go deeper

for a junior

Know that a screening stage is not a function: the same item can be judged differently on different runs, so a single pass is one observation rather than a property of the stage.

for a middle

Explain what has to be pinned for the result to mean anything: which stage, which version, what text it actually judged, how many trials and what comparable items did.

for a senior

Show the judgment — weigh the pass rate against retry cost in that path, and separate a construction that survives rewording from one string that will not outlive the next rebuild.

for a principal

Be ready to say what your programme does with unreproducible-but-plausible results, and how reports are pinned so a successor can tell a fixed issue from a stale test.

## The question behind the question An interviewer asking this is not asking for a threshold. They are asking whether you know what a single probabilistic result is evidence of, and whether you can tell a finding about a *mechanism* from a finding about *one string against one build*. ## What the one pass actually proves It proves that on one occasion a stage emitted a passing verdict for a submitted item. That is a narrower statement than it feels like, and every word in it carries weight: **on one occasion** (sampling exists), **a stage** (which one?), **for a submitted item** (was it the item the stage read?). ## Three things to establish before it is a finding **1. Which stage decided.** Screening paths are usually stacked — a cheap label-only classifier in front, a generative stage behind, sometimes an output-side stage as well. If the platform records a final outcome only, the report cannot say whether the passing item cleared a stage that reads text or one that does not, and those are entirely different claims about entirely different components. A label-only stage additionally produces no account of its own decision, so there is no rationale to cite even when the outcome is recorded. This is the concrete reason such findings stall: the evidence the report needs was never produced. **2. Whether the stages read what you submitted.** Normalisation, extraction and truncation sit between submission and judgment. If the run cannot show that the judged text matched what was written, an inconsistent result may be a difference in what arrived rather than a difference in how it was judged. **3. Whether one in five beats noise.** A stage with any sampling has a base rate of variation, and comparable items that were *not* written at the screen also pass sometimes. Without a control arm, one in five is a number with no denominator behind it. Five runs is also a very small sample: the confidence interval around one-fifth on five trials is wide enough to be nearly useless, which is a fact about statistics and not about the screen. ## Then the judgment call Once it stands up, rate is weighed against **retry cost**, not against a general sense of reliability. A path where an item can simply be resubmitted converts a low per-attempt rate into eventual passage, so a fifth is material. A path where an attempt costs an account, a delay or a review escalation is a different economic picture entirely. The other axis is **breadth**. A construction that keeps working when the wording is changed is evidence about a mechanism; one that dies the moment a synonym is substituted is evidence about a string, and strings do not survive a rebuild. The first is worth filing and worth writing up; the second is worth a note. ## Reproduction across versions Stages get rebuilt. If the item now fails and the stage has changed since the report, the failure is not evidence the construction never worked — it is evidence the deployment moved. Reports that do not pin the stage version at the time of the run become unfalsifiable within weeks, and the person who inherits the report cannot tell a fixed issue from a stale test. Pinning what was observed — which stage, which version, which submitted text, how many trials, what the control arm did — is what makes the difference between a claim someone can act on and a claim nobody can settle. ## The failure mode to avoid in both directions Over-claiming: `the screen is bypassable` from one pass. Under-claiming: `it did not reproduce, so there is nothing here` after four failures in five, on a path with free retries. Both are the same error — reading a rate as a boolean.

  • With a label-only stage in front and a generative stage behind, what stops you attributing the pass?
    If the platform records only a final outcome, nothing in the evidence says which stage emitted it. The label-only stage also produces no account of its own decision, only a score, so even where its outcome is recorded there is no reasoning to cite. Without per-stage decisions the report describes an outcome and cannot name a component.
  • The item fails today and the stage has been rebuilt since. What can you conclude?
    That the current build behaves differently, and nothing more. It is not evidence the original observation was wrong, and it is not evidence anything was deliberately addressed. This is why a report pins the stage version, the submitted text and the trial count at the time of the run — without those it becomes unsettleable rather than resolved.
  • How do you tell a mechanism result from a one-string result?
    Vary the wording while keeping the property you believe is doing the work, and see whether the effect survives. If it does across several rewrites, the claim is about a class of construction. If it dies on the first substitution, the claim is about one string against one build, which will not survive the next rebuild and should be written up as much weaker.

saying these in an interview costs you the question

  • Calls one pass in five a bypass without qualification
  • Discards a low rate on a path with free retries
  • Cannot say which stage emitted the passing verdict
  • Ignores that the judged text may differ from what was submitted
  • Treats five trials as a meaningful sample
  • Reads a failure after a rebuild as disproof of the original run

context