A failing test saved the generative step's input and output but nothing about the call itself. What can an investigation not conclude?
answer
- Four facts describe the call
- Input and output are only the ends
- Which wording, which model identifier
- Without settings, a repeat proves nothing
basics
~20 sWithout the supplied material, the instruction version, the model identifier and the generation settings, an investigation cannot say which wording or model produced the failure, whether a repeat is the same call, or whether a later edit fixed anything.
solid answer
~50 sInput and output describe the ends of the attempt, not the attempt. Four facts describe the call, and every one of them can change without anyone touching product logic: the material the product assembled and supplied, the instruction text with its version, the model identifier, and the generation settings. Drop the supplied material and you cannot tell a bad answer from a reasonable answer to badly assembled content. Drop the instruction version and no later edit can be credited with the fix, because you never knew which wording failed. Drop the model identifier and you cannot say the release was exercising the model it ships against. Drop the generation settings and a repeat is a different call, so a non-reproduction proves nothing. The outcome of missing them is not a wrong conclusion but an unfalsifiable one: the failure is closed as unexplainable and returns unchanged.
code
pseudocode · 12 linesREQUIRED_CALL_FACTS = [suppliedMaterial, instructionVersion,
modelIdentifier, generationSettings]
on write_record(record):
missing = [f for f in REQUIRED_CALL_FACTS if record[f] is absent]
if missing is not empty:
fail_build("evidence record incomplete", missing)
store.put(record)
on investigate(record):
if record.modelIdentifier != record.servedByModel:
note("called and served identifiers differ")go deeper
Know that a saved input and a saved answer are not a complete record of what happened, and that the wording and settings the product used are separate facts somebody has to write down.
Name the four call facts and what each one is for: the supplied material, the instruction text and its version, the model identifier, and the generation settings that decide whether two attempts are comparable at all.
Demonstrate the diagnostic consequence of each absence, and show how you keep the recorded facts honest by reading them from the call itself and failing loudly when the evidence writer produces a record with a hole in it.
Argue for where these facts are guaranteed rather than requested: a shared call path that records them by construction, versus per-team discipline that decays, and what you accept in exchange for the coupling that creates.
## Input and output describe the ends, not the call An investigator arriving at a failed test of a feature that answers from a generative step has exactly one question: what actually happened? The input and the output describe the two ends of that attempt. They say nothing about the attempt itself. Four facts do, and each of them can change underneath a team without anybody editing a line of the product's own logic. 1. **The material supplied to the generative step** — whatever the product retrieved, assembled or attached on that attempt. It is assembled at run time, so it is different on every attempt by construction. 2. **The instruction text and its version.** The wording is an artefact of the product like any other, and it is edited far more often than product code because editing it is cheap and feels safe. 3. **The model identifier.** Which model answered, at which version. A product that follows a moving alias is being changed by someone else on a schedule it does not control. 4. **The generation settings.** How much variation is permitted, any ceiling on output length, any stop condition. These decide whether two attempts are even comparable. ## What each absence costs | Fact missing from the record | What the investigation can no longer conclude | |---|---| | The supplied material | Whether the step answered badly, or answered reasonably from material the product assembled wrongly | | The instruction text version | Which wording produced the failure, so no later edit can be credited with fixing it and no rollback can be aimed | | The model identifier | Whether the release was even exercising the model it ships against, or whether the failure predates a change of model | | The generation settings | Whether a repeat is the same call at all, so a non-reproduction proves nothing | Read that table the other way and it becomes a design rule: **record whatever, if absent, would let someone answer "we cannot tell" and close the investigation.** ## The failure mode this prevents The characteristic bad outcome is not a wrong conclusion. It is an **unfalsifiable** one. Somebody says the generative step was having a bad day; nobody can show otherwise, because the evidence does not contain the facts that would contradict it. The failure is closed as unreproducible, the same shape of failure returns three weeks later, and the second investigation starts from the same standing start as the first. The second bad outcome is a **false all-clear**. A wording is edited, the test passes, and the change is credited with the fix. If the record never said which wording ran when it failed, the pass is equally consistent with the failing wording having been fine and the attempt simply having gone the other way. You have not fixed anything; you have observed a different sample. ## Recording them where they cannot fall out of step Facts recorded from a second source drift away from the facts that were used. Two habits keep them honest: - **Read the value from the same object the call used, at call time.** Take the instruction version from the artefact that was passed to the step, not from a configuration file read separately by the reporting code. Take the model identifier from what the call was actually made with, and where the response reports back which model served it, record that too and record the mismatch when they differ. - **Make the record fail loudly when a fact is absent.** An evidence writer that quietly omits an unset field produces records that look complete. Writing an explicit absent marker, or refusing to accept an evidence record with a hole in it, converts a silent gap into a build-time complaint while somebody can still fix it. ## What this still does not settle Recording the four facts makes an investigation possible; it does not perform it. The record will not tell you whether the failure is a defect in the product or a limit of what the feature can do, and it will not tell you whether the case should be repeated or the assertion loosened. Those are separate decisions, taken by people, with this material in front of them. What the record guarantees is narrower and more useful than a verdict: that when those decisions are taken, they are taken about a specific call that specifically happened, rather than about a reconstruction assembled from memory a fortnight later.
- The model identifier was recorded but the generation settings were not. What is still open?You know which model answered but not whether any later attempt is the same call. How much variation was permitted, any ceiling on output length and any stop condition all shape what comes back, so a repeat under different settings that behaves cannot exonerate the failing attempt, and a repeat that misbehaves cannot confirm it.
- How do you stop the recorded instruction version from falling out of step with what actually ran?Read it from the same object the call was made with, at call time, rather than from a configuration source the reporting code consults separately. Where the response reports which model served it, record that alongside the one you asked for and flag a mismatch, because a moving alias changes what answered without changing anything you wrote.
saying these in an interview costs you the question
- Treats the model identifier as enough without the generation settings
- Records the instruction text but not which version of it ran
- Reads the instruction version from configuration rather than from the call
- Assumes a repeat that passes proves the defect is gone
- Lets the evidence writer skip fields that happen to be unset