skip to content

Release Judgement

The calls a team makes about a feature that is only usually right: what residual risk to state, what failure rate is acceptable, and which layer owns a complaint. The judgement is what is probed.

on this pageshow

questions

8

Which layers can own a wrong answer from an AI-powered product feature, and how do you decide which one it is?

level: middleimportance: must knowfreq 62%

answer

  1. a symptom, not yet a diagnosis
  2. most of the feature is ordinary code
  3. five layers, only one probabilistic
  4. read what was actually sent
  5. change one layer, re-measure the rate

basics

~20 s

Five layers can own it: the ordinary product code around the generative step, the instruction text sent to it, the material supplied with it, the configuration, and the generative model's own capability ceiling. Attribute only after evidence separates them.

solid answer

~50 s

A wrong answer is a symptom, not a diagnosis. Five layers can produce it: - **Product code** around the generative step: a long reply cut, a cached answer served, an empty result rendered as a confident one. - **Instruction text** the feature sends: ambiguous wording, a missing rule, two directions that contradict. - **Supplied material**: the fact the answer needs was never assembled, arrived out of date, or was cut to fit a size limit. - **Configuration**: a variant, a randomness setting or a model version differing from what was tested. - **The generative model's capability ceiling**: the task is beyond what it does reliably at any wording. Decide on evidence. Reconstruct the exact call, read what was actually sent, repeat the identical input enough times to get a rate, then change one layer at a time. The ticket goes to the layer whose change moves that rate.

code

pseudocode · 18 lines
pseudocode
recorded = reconstructCall(complaint)
// recorded: userInput, instructionText, suppliedMaterial, settings, returnedText

if recorded.returnedText acceptable AND complaint.shownAnswer != recorded.returnedText:
    owner = "product code"          // parsing, caching, truncation, rendering

else if neededFact not present in recorded.suppliedMaterial:
    owner = "supplied material"     // never assembled, stale, or cut to fit

else if recorded.settings != testedSettings:
    owner = "configuration"

else:
    baseRate = failureRate(repeat(recorded, 20))
    editedRate = failureRate(repeat(withEditedInstructions(recorded), 20))
    if editedRate << baseRate: owner = "instruction text"
    else if baseRate < 0.1:    owner = "variation, no single layer"
    else:                      owner = "capability ceiling, record it"

go deeper

for a junior

Be ready to name the parts of an AI-powered feature: code that gathers material, code that assembles instructions, the generative step, and code that renders the reply. Knowing that four of the four surrounding parts are ordinary code is the whole starting point.

for a middle

Explain the mechanics of separating the layers: reconstruct the exact call, compare what came back with what was shown, read the material as sent, and change one thing at a time. Interviewers expect you to reach for evidence before opinion.

for a senior

Show the production judgment: measure rates over fixed inputs rather than trusting single reproductions, decide which layer to correct first when two look guilty, and route the resulting ticket to the owner whose change actually moved the rate.

for a principal

Own the routing itself. A single queue that attributes every wrong answer to the model integration hides four other defect classes, so argue for triage that splits them, and for the recording that lets the team see which layer is really generating its complaint volume.

## A wrong answer is a symptom, not a diagnosis An AI-powered product feature is ordinary software with one probabilistic component inside it. A user request arrives through normal handling; the feature gathers supporting material; it assembles instruction text; it calls a generative model; it parses, validates and renders whatever comes back. Four of those five stages are deterministic code the team wrote and can reason about exactly. Only one is not. When a complaint arrives, the probabilistic stage is the most memorable suspect and frequently the least likely culprit, precisely because it is the part nobody on the team wrote. **Attribution** is naming, on evidence, which stage produced the failure, before anyone writes a fix. It matters because each layer has a different owner, a different fix cost and a different confidence that the fix will hold. Product code is cheap to fix and the fix is permanent. Instruction text is cheap to change and the change is fragile. What material reaches the generative step is a pipeline concern with its own tests. And *the generative model cannot do this reliably* is not a fix at all: it is a constraint the surrounding product has to be designed around. ## The five layers, and the evidence that convicts each | Layer | What it looks like | Evidence that convicts it | | --- | --- | --- | | **Product code around the generative step** | The answer shown differs from what came back; a cached answer is served; a long reply is cut; an error path renders an empty result as a confident one | Replaying the recorded call yields an acceptable output while the rendered answer stays wrong | | **Instruction text the feature sends** | Ambiguous wording, a missing rule the answer depends on, or two directions that contradict; the failure repeats across many inputs of the same shape | One targeted edit to the wording flips the outcome repeatedly on the same fixed inputs | | **Material supplied with the request** | The fact the answer needs was never assembled, arrived out of date, or was cut to fit a size limit | Reading the material exactly as sent shows the fact is absent, and supplying it makes the answer correct | | **Configuration** | A variant, a randomness setting, a size limit or a model version differs from the one that was tested | The failing and passing calls differ only in a setting, and restoring the tested value restores the outcome | | **The generative model's capability ceiling** | Complete material, careful wording and tested settings still fail at a stable rate across a fixed set of inputs | Nothing available to the team moves the measured rate | ## Working through it 1. **Reconstruct the call.** Pull the literal user input, the instruction text as assembled, the material as sent, the settings in force, and the text that came back. Without those five, everything after this is guesswork. 2. **Check the deterministic shell first.** Compare what came back with what the user saw. A mismatch ends the investigation in ordinary product code, and on a mature feature that is the most common outcome. 3. **Read what was actually sent, not what the code was meant to send.** Missing, stale or truncated material is visible here without calling the generative model again, which makes it the cheapest evidence in the whole investigation. 4. **Repeat the identical call.** One wrong answer is one sample. A failure that appears once in twenty identical repeats is variation rather than a defect with a single owner; a failure that appears in every repeat is systematic and worth chasing. 5. **Change one layer at a time and re-measure.** The layer whose change moves the failure rate is the layer that owns the ticket. If no change moves it, a capability ceiling is the remaining explanation, and that conclusion has to be earned rather than assumed. ## Who gets the ticket - Product-code failures go to whoever owns that code, exactly like any other defect; nothing about the feature being AI-powered changes that. - Instruction and material failures go to the team that owns the assembly of the request. Often those are the same engineers, but the fix and the test that protects it are different in kind. - Configuration failures usually travel with a second finding about how an untested setting reached users at all, which is frequently the more valuable half. - A ceiling gets no fix ticket. It gets a recorded limit and a product decision about what the feature does when it hits that limit. The common anti-pattern is a single queue where every wrong answer is filed against whoever integrated the generative model. That queue mixes five defect classes with five different fixes, and it hides the fact that most of its contents are ordinary bugs in ordinary code. ## What one reproduction cannot tell you Because the generative step can return different text for the same input, evidence on this feature is statistical in a way it is not elsewhere in the product. Attribution therefore rests on rates measured over fixed inputs, not on single outcomes. An experiment that *worked when I tried it* has told you almost nothing, and a fix accepted on the strength of one passing attempt is a fix nobody has tested.

  • Why is one reproduction of the wrong answer not enough to attribute it to the generative model?
    The generative step can return different text for identical input, so a single wrong answer is one sample from a distribution. It could be the unlucky end of an otherwise acceptable rate. Repeat the identical call enough times to compare rates before and after each change; a conclusion drawn from one attempt cannot survive the next attempt disagreeing with it.
  • Two layers look guilty at once: the wording is vague and the supplied material was out of date. Which do you address first?
    The material, because it is deterministic and verifiable. Correct the assembly, re-measure on the same fixed inputs, and see what is left. Fixing wording first leaves you unable to say which change helped, and a wording change that compensates for bad material is a fix that quietly breaks when the material is later corrected.
  • What does a team lose by filing every wrong answer against whoever owns the generative-model integration?
    It loses the shape of its own defect population. One queue then mixes rendering bugs, assembly bugs, untested settings and genuine ceilings, which have different owners, fix costs and fix durability. The visible effect is a large backlog attributed to the model while most of it is ordinary code that would be fixed in an afternoon if it were routed correctly.

Treat it like a wrong figure on a printed invoice: the arithmetic engine is the last thing you suspect, after the data fed in, the template and the printer.

saying these in an interview costs you the question

  • Blames the generative model before reading what was actually sent
  • Treats one wrong answer as proof of a systematic defect
  • Rewrites the instruction wording and ships without re-measuring
  • Files every wrong answer against whoever integrated the model
  • Claims the feature cannot be investigated because output varies
open as a page

For a feature whose answers come from a generative step, why does a test suite pass rate misstate readiness, and what replaces it?

level: middleimportance: must knowfreq 58%

basics

~20 s

A pass rate describes one judged sample of varying output, not the feature itself: re-judge the same cases and the number moves. Report a measured residual error rate per user action, with its sample size, judging rule and split by consequence.

open as a page

The same build of an AI-powered feature answers one user correctly and another wrongly. What differences do you rule out first?

level: middleimportance: should knowfreq 45%

basics

~20 s

Rule out deterministic differences before blaming any layer: the literal inputs each user sent, the material each account's request assembled, the variant and settings served, and the client each used. Only then consider run-to-run variation.

open as a page

An AI-powered feature returned a wrong answer. Which cheap experiments separate the possible causes, and in what order do you run them?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Run the cheap deterministic checks first: replay the recorded call, compare the shown answer with the returned one, and read the instructions and material exactly as sent. Only then vary one layer at a time and re-measure the failure rate.

open as a page

When no change your team can make fixes an AI-powered feature's wrong answers, what must that conclusion rest on, and what is recorded instead of a fix?

level: seniorimportance: should knowfreq 38%

basics

~20 s

It rests on evidence, not exhaustion: a measured failure rate on fixed inputs, and every cheaper layer changed and re-measured one at a time. In place of a fix, record the limit, its conditions, its containment, and when to re-check.

open as a page

Not every wrong answer from a generative product feature costs the same. How do you separate cheap misses from expensive ones?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Classify each wrong answer by what it costs rather than how often it happens: whether the user notices at the time, whether it can be undone, and whether the feature acted or only proposed. Then set a separate tolerance for each class.

open as a page

You sign off a feature whose answers come from a generative step and will sometimes be wrong. What observation reverses that decision, and who watches?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Wrong answers are forecast, so one bad output reverses nothing. Write a rate on an observable signal — edits, discards, repeat attempts, escalations — with a review level, a revert level, a time window, a named person watching, and a fallback that already exists.

open as a page

A product feature's generated answers are wrong in 4% of attempts and will not reach zero. How do you make the ship call?

level: principalimportance: should knowfreq 42%

basics

~20 s

Compare against the process the feature replaces, not against zero. Establish whether the figure is trustworthy, what each wrong answer costs, whether errors are detectable by the user and by the team, and what reverting takes. Then record the number, its conditions and who accepted it.

open as a page