Why should a model-judged verdict never be the only thing that fails a build, and what deterministic check sits behind it?
answer
- What must a failing build mean on re-run?
- Two layers, two different powers
- Failure always wins over judgement
- Confirmed findings become explicit assertions
basics
~20 sA build gate must be reproducible, reviewable and stable in its wording; a judged verdict is none of these. Let explicit assertions on statuses, amounts, identifiers, required fields and forbidden content fail the build, and report judged findings as signals.
solid answer
~40 sA failing build has to mean something on a re-run, and a judged verdict is generated rather than compared, so a re-run may clear it or invent a new one. Its explanation is reworded each time, so identical problems cannot be grouped, and gating on it adds an external dependency whose outage stops merges. The working arrangement keeps both layers with different powers: **explicit assertions decide the build** — status and outcome, amounts, identifiers, required fields, invariants, forbidden content such as placeholder text or another customer's data — while **judged findings are reported** as annotations or a review queue. A judged pass must never clear a deterministic failure. When the judged layer finds a real defect and a human confirms it, promote that defect to an explicit assertion so the gating layer ratchets upward.
code
pseudocode · 12 linesdeterministic_failures = run_explicit_assertions(case)
judged_findings = run_judged_checks(case)
if deterministic_failures is not empty:
fail_build(deterministic_failures) # the gate, and the only gate
# a judged pass never clears the above; it is reported either way
publish_signal(judged_findings) # annotation + review queue
for finding in judged_findings:
if confirmed_by_human(finding) and seen_in_previous_runs(finding):
open_task("write an explicit assertion for", finding)go deeper
Recall the rule of thumb: a model's opinion may be reported, but the thing that stops a release compares values a person wrote down. Know that a judged pass never overrides a check that failed.
Explain the split and what belongs on each side: statuses, amounts, identifiers, required fields, invariants and forbidden content asserted explicitly; open-ended quality judged and reported as a signal.
Show the operating consequences of gating on an unreproducible verdict — re-runs that clear real failures, failure text that cannot be grouped, an external dependency in the merge path — and describe the promotion path from confirmed finding to explicit assertion.
Own the policy and its economics: what authority an advisory layer is allowed, who reviews its findings, how promotion is funded so the gating layer ratchets upward, and how a team is stopped from quietly making an advisory check a required one.
## What a build gate has to be A check that can stop a release is held to three properties that have nothing to do with how clever it is. It must be **reproducible**, so a red result means the same thing when someone re-runs it. It must be **reviewable**, so a disputed failure can be examined rather than argued about. And it must be **legible**, so its failure text says the same thing every time and can be grouped, counted and tracked. A judged verdict has none of the three by default. The deciding step samples, so a re-run may clear a real failure or invent a new one. Its reasoning is regenerated in new words on every call, so two occurrences of the same problem read as two different problems. And the standard it judges against is prose, which means people disagree about what the failure even asserted. On top of that the build acquires a new runtime dependency with its own latency, cost and availability: an unrelated outage now stops merges, and the outage is not in your product. There is a further, quieter risk. A judged verdict that can fail a build will, sooner or later, be argued about under time pressure, and the way that argument ends is a re-run. Once a team learns that pressing the button again clears a red build, the deterministic failures standing next to it get the same treatment. That is the real cost: an unreliable gate does not just fail to help, it erodes the gates that were working. ## The split that works Keep both layers in the case and give them different powers. - **The deterministic layer decides the build.** Status and outcome, amounts and totals, identifiers, required fields present and correctly typed, values in range, invariants holding, forbidden content absent — placeholder text, an internal stack trace, another customer's data. All of it explicit, all of it exact. - **The judged layer produces signals.** Its findings are attached to the run as annotations, posted to a review queue, or collected into a report a human reads. It never turns a green run red on its own. - **A judged pass never clears a deterministic failure.** The dangerous direction is not the judged layer complaining too much, it is a confident judged pass being read as reassurance over a check that actually failed. Failure must always win over judgement. - **Repeated, confirmed findings get promoted.** When the judged layer finds a real defect and a human confirms it, write an explicit assertion for that defect and let it join the deterministic layer permanently. That last point is the mechanism that makes the whole arrangement pay. The judged layer is a frontier scout, not a fence. Every real thing it finds should become a deterministic check, so the suite's gating power ratchets upward while the judged layer keeps watching only the part of the surface no assertion covers yet. | Concern | Deterministic layer | Judged layer | | --- | --- | --- | | May fail the build | Yes, always | No, reports only | | Reproducible | Guaranteed | Not guaranteed | | Failure text | Fixed and groupable | Regenerated each run | | Depends on an external service | No | Yes: latency, cost, availability | | Role over time | Accumulates confirmed checks | Scouts the uncovered frontier | ## Objections you should be able to answer *"Then the judged layer has no teeth."* It has the teeth its evidence supports. A judged finding that recurs across runs on the same output is strong; a single flapping verdict is not. If you want teeth, spend them on promotion: turn the confirmed finding into an assertion that does gate. *"We can require agreement across several judgements."* Repeated judgement genuinely reduces flapping and is worth doing, but it narrows the variance rather than removing it, and it multiplies latency and cost on every build. It makes a better signal; it does not make a gate. *"Our judged check has been right for months."* Being right is not the same as being reproducible. The property a gate needs is that the same input gives the same answer, and no run of good luck supplies it. Meanwhile the judge can be updated at any time without a commit in your repository. The summary a strong candidate gives: keep the judged verdict as a high-reach, low-authority signal sitting beside a low-reach, high-authority deterministic core, and run a deliberate pipeline from the first to the second. That way the suite gains the reach without spending the one property that makes a failing build worth believing.
- What should happen when the judged layer keeps finding the same real defect release after release?Promote it. Once a human has confirmed the finding and it has recurred, write an explicit assertion that catches that specific defect and let it join the gating layer permanently. The judged layer then goes back to watching the part of the surface no assertion covers. Leaving a confirmed, repeatable defect in an advisory report is the failure mode: the suite never gains the check it already earned.
- Would requiring three agreeing judgements make a judged verdict safe to gate on?It makes a better signal, not a gate. Repeated judgement narrows the variance but does not remove it, and it multiplies latency and cost on every build while leaving the external dependency in the critical path. The property a gate needs is that the same input always gives the same answer, and no amount of repetition supplies that. Spend the effort on promoting confirmed findings instead.
- Which direction of judged error is the more dangerous next to a deterministic layer?A judged pass being read as reassurance. A noisy judged layer costs attention and is visible, but a confident pass over output that a deterministic check actually failed, or over a surface nobody else examines, quietly removes a check people believe exists. Failure must always win over judgement, and a broad judged pass must never be recorded as coverage.
saying these in an interview costs you the question
- Letting an unreproducible verdict turn a green run red
- Reading a confident judged pass as reassurance over a real failure
- Re-running the build until the judged verdict clears
- Leaving confirmed judged findings as findings forever
- Adding a build dependency on an external service's availability