skip to content

Why can a check that asks a model to judge a page's error text pass one run and fail the next on unchanged output?

level: juniorimportance: should knowfreq 38%

answer

  1. The deciding step generates, not compares
  2. Which outputs sit near the line?
  3. The written standard's wording matters
  4. The judge itself can change underneath you

basics

~20 s

A judged step generates its answer by sampling, so identical input can produce different verdicts. Borderline outputs flip most, wording changes in the standard move verdicts, and the model doing the judging can be updated without a commit anywhere.

solid answer

~40 s

The deciding step is a generation, not a comparison, so the same input can legitimately yield a different answer twice. Outputs that sit near the line the written standard draws flip most often, which is why a flapping judged check usually means the case is borderline against a vague standard. Rewording or reordering that standard also moves verdicts, and if the judging step receives the whole page it also receives timestamps and generated identifiers that change every run. Finally, the model doing the judging is a dependency that can be updated underneath a suite that did not change. The consequence matters more than the cause: failures are no longer reproducible, failure text is regenerated in new words each time, and a red that clears on re-run teaches everyone to re-run.

code

pseudocode · 14 lines
pseudocode
# make the judged step's inputs part of the record
inputs = {
    standard_version: "refund-message-v4",
    standard_text:    load("standards/refund-message-v4.txt"),
    output:           extract_paragraph(page, "error-banner"),
    randomness:       "lowest setting available"
}

verdicts = repeat 3 times: judge(inputs)

if not all_equal(verdicts):
    report("unstable-verdict", inputs, verdicts)   # the disagreement IS the finding
else:
    report(verdicts[0], inputs)

go deeper

for a junior

Be ready to say that a model-judged step generates an answer rather than comparing values, so the same input can give different answers. Knowing that much stops you from hunting for a race in the product first.

for a middle

Explain the mechanics behind the variance: sampling, borderline inputs near a vague standard, sensitivity to the standard's wording, volatile values leaking into the input, and updates to the judge itself.

for a senior

Demonstrate what the variance costs an operating suite: failures you cannot reproduce or bisect, failure text that cannot be grouped, and a team learning to re-run. Show the containment: versioned standards, narrow inputs, repeated judgement, recorded artefacts.

for a principal

Own the trade explicitly: judged checks buy reach on output nothing else can examine and pay in reproducibility. Decide where that trade is acceptable and how the team is prevented from spending it on checks that had a deterministic option.

## Why the same input does not give the same verdict A check whose outcome is decided by a model is not a comparison, it is a generation. The judging step produces its answer by sampling, so two calls with byte-identical inputs can legitimately return different answers. That is a property of the mechanism, not a bug in your case, and it is the first thing to say when a judged check flickers on unchanged output. Four further sources of variance sit on top of the sampling itself. - **Borderline inputs.** Most outputs are obviously fine or obviously broken and get a stable verdict. A minority sit near the line the standard draws — an error message that is *almost* actionable — and those are the ones that flip. A flapping judged check is usually telling you that the case you chose is borderline against the standard you wrote. - **Instruction sensitivity.** Rewording the standard, reordering its sentences, or adding an example moves verdicts on those borderline cases. Two instructions a human reads as identical do not necessarily judge identically. - **Drifting inputs you did not mean to send.** If the judged step receives the whole page or the whole response, it also receives timestamps, generated identifiers, session-specific text and whatever else changes run to run. Any of that can tip a borderline judgement. - **The judge changing underneath you.** The model behind the check is a dependency like any other. When it is updated, verdicts move on outputs that did not, and nothing in your repository records the change. ## What that does to a suite A test suite is valuable mostly because a red result means something. Non-determinism in the deciding step attacks exactly that. 1. **Failures stop being reproducible.** Re-running does not confirm a failure, it re-rolls it. 2. **Bisecting stops working.** Walking back through changes needs a stable decision function; here the decision itself moves while you search. 3. **Failure text stops being stable.** A judged failure explains itself in fresh words each time, so identical problems cannot be grouped, counted or tracked over time. 4. **A red that goes green on a re-run trains the team to re-run.** Once people learn that pressing the button again clears the build, the deterministic failures sitting next to it get the same treatment. 5. **History becomes uncomparable.** A pass rate that moves because the judge moved looks exactly like a pass rate that moved because the product moved. | Property a suite relies on | Explicit comparison | Judged verdict | | --- | --- | --- | | Same input, same outcome | Guaranteed by construction | Not guaranteed at all | | Failure text | Fixed, mechanical, groupable | Regenerated in new words each time | | Reproducing a failure | Re-run confirms it | Re-run may silently clear it | | Dependency that can shift | The code under test only | The judge, and its update schedule | ## Making a judged check as stable as it can be You cannot make sampling deterministic, but you can shrink what varies and make the variance visible instead of silent. - **Version the standard.** The judging instruction is a test asset. Keep it in the repository, give it a version, and record which version produced a verdict. - **Send the narrowest input.** Extract the one paragraph under examination rather than the whole page, and strip the values that change every run before the judging step ever sees them. - **Ask one narrow question.** Narrow questions have fewer borderline cases, and fewer borderline cases is most of the stability you can win. - **Use the least random setting available**, and record it next to the verdict. - **Judge more than once when it matters.** If three judgements of the same input disagree, that disagreement is the finding: report it as an unstable verdict rather than letting whichever answer arrived first stand as the result. - **Record the inputs and the stated reasoning as artefacts** of the run, so a disputed verdict can be examined afterwards instead of re-litigated by re-running. - **Re-check the standard when the judge changes.** Treat a judge update the way you would treat a dependency upgrade: run it against known outputs before believing new verdicts. The honest summary for an interview: a judged check trades reproducibility for reach. That trade is sometimes worth making on output nothing else can examine, but it is never invisible, and a team that does not plan for the variance discovers it as an unexplainable flapping build instead.

  • Three judgements of the same output disagree. What should the check report?
    Report the disagreement itself as an unstable verdict, with all three answers and the exact inputs attached. Letting whichever answer arrived first stand as the outcome hides the most useful signal you have, which is that this output sits on the line the standard draws. The repair is usually to sharpen the standard or narrow the question, not to add more retries.
  • Why does sending the whole page to the judging step make verdicts less stable?
    Because the page carries values that change on every run: timestamps, generated identifiers, session-specific text, unrelated components. All of that becomes part of the input being judged, so the judging step sees a different input each run even when the paragraph under examination is identical. Extract the narrow fragment you actually care about and strip volatile values first.

saying these in an interview costs you the question

  • Assuming the flapping must be a race in the product
  • Re-running until the judged verdict comes back green
  • Keeping the judging standard only in someone's head
  • Feeding the whole page into the judging step
  • Treating the judge as a fixed, unchanging dependency