skip to content

In a generative product feature, which evidence shows a varying test result comes from the feature rather than the test?

level: seniorimportance: should knowfreq 50%

answer

  1. pin one side and repeat
  2. which assertion actually moved?
  3. alone, then inside the whole suite
  4. misses cluster on inputs, or on machines
  5. did unrelated cases go red too?

basics

~20 s

Pin the generative step to one fixed recorded answer and repeat: if the result still moves, the test is the source. Then read which assertion moved - generated wording varies by design, a stored identifier does not.

solid answer

~50 s

Gather evidence before changing anything, because the two have different owners and each repair destroys the other diagnosis. First pin one side: replay a single fixed answer for every repeat, and if the result still moves, the variation lives in the test or the code around the generative step. Second, read *which* assertion moved - wording, ordering and phrasing are what the generative step is free to change; an identifier, a stored row or a status is not. Third, run the case alone and then inside the full suite, because variation that appears only under the suite points at ordering, shared data or contention. Fourth, see what the misses correlate with: variation belonging to the feature tracks the input and holds a stable rate, while unreliability belonging to the test tracks parallelism, load and which machine ran it.

code

pseudocode · 10 lines
pseudocode
# control experiment: same input, generative step pinned to one recorded answer
fixed = recorded_answer_for(input)

for i in 1..10:
    output = feature.answer(input, generative_step = replay(fixed))
    record(i, meets_criteria(output), first_unmet_criterion(output))

# result still moves  -> variation lives in the test or the surrounding code
# result now constant -> variation is the product's generative step
# put the case back to its normal wiring afterwards

go deeper

for a junior

Know that a result can move for two different reasons - the feature genuinely varies, or the test around it is unreliable - and that the repair is completely different in each case.

for a middle

Be able to describe the control experiment: pin the generative step to one fixed recorded answer, repeat the case, and see whether the result still moves.

for a senior

Show a diagnosis order rather than a hunch: pin one side, read which assertion moved, run alone versus inside the suite, and check what the misses correlate with before changing anything.

for a principal

Own how the call is made: what evidence the team requires before a case is labelled feature variance, and who reviews that label, so it does not become a way of explaining reds away.

## Two sources that produce the same symptom A test case exercising a product feature whose answer comes from a generative step sits on top of two independent sources of variation. The first is the feature: identical input, different output, by design. The second is everything a suite always had - timing, ordering, shared data, concurrency, an environment that moves underneath it. Both produce the same surface symptom: the case passed yesterday and failed today with nothing committed. They are completely different problems. Variation that belongs to the feature is not a defect; it is the thing a rate rule exists to measure, and the response is a decision about that rate. Variation introduced by the test is a defect in your own code, and the response is to fix it. Guessing between them is how teams end up loosening a rule to accommodate a shared-state bug, or rewriting a perfectly good test because a feature was always going to vary. ## The evidence, gathered before anything is changed 1. **Pin one side and repeat.** Replace the generative step with a single fixed recorded answer for every repeat and run the case ten times. If the result still moves, the varying input is gone and the variation must live in the test or the code around it. If it goes constant, the generative step is the source. This is a control experiment, not a change to the case - put the case back afterwards. 2. **Read which assertion moved.** Free text - wording, ordering, phrasing, which of several acceptable facts is mentioned first - is exactly what a generative step is free to change between attempts. An identifier, a stored row, a status, a count of created records, the presence of a required field: nothing about a generative step gives it permission to change those. A miss on the deterministic side is not the feature varying. 3. **Run it alone, then inside the whole suite.** A case that is stable in isolation and unstable under the full suite is telling you about ordering, leftover data or contention between tests, none of which the feature can cause. 4. **Check what the misses correlate with.** Variation that belongs to the feature tracks the **input**: the same hard input misses more often, and the rate is roughly stable over time. Unreliability that belongs to the test tracks the **machine**: parallelism, load, time of day, which worker picked the case up. 5. **Check whether neighbours moved too.** If cases that never touch the generative step went red on the same build, the cause is shared - environment, credentials, data, a dependency - and reading anything into one case's rate is premature. | What you observe | Most likely source | | --- | --- | | result still moves with the generative step pinned to one answer | the test or the surrounding code | | the failing assertion is on generated wording | the feature | | the failing assertion is on a stored identifier | the test or the surrounding code | | stable alone, unstable inside the full suite | the test | | misses concentrate on one input, rate otherwise stable | the feature | | misses track parallelism or which machine ran it | the test | | unrelated cases went red on the same build | something shared, neither one | ## Why the order matters Both repairs are destructive to the other diagnosis, which is why the evidence comes before the edit. Loosening the criteria "because the feature varies" hides a shared-state defect permanently, and the loosened case will now pass through a real regression as well. Rewriting the test "because it is unreliable" throws away a genuine measurement of a falling success rate - and the rewritten case will vary exactly as much as the old one, because the variation was never in the code you rewrote. Two habits make the evidence available at the moment you need it: - Keep the per-repeat outcomes, not only the final result. You cannot see that a rate fell from five-in-five to three-in-five if only the final mark was kept. - Keep assertions over deterministic values separate from assertions over generated text, so which side moved is visible in the failure itself instead of needing an investigation to recover. One more thing worth saying out loud: **the two can be present at once.** A case can be genuinely variable *and* sit on a leaked fixture. The pin-one-side experiment is what stops you from stopping at the first explanation that fits.

  • Every case in the suite went red on the same build, including ones that never touch the generative step. What does that tell you?
    That the cause is shared rather than per-case: environment, data, credentials or a dependency that everything leans on. A feature's own variation does not synchronise across unrelated cases. Confirm it by checking the cases with no generative step in them, fix the shared cause, and only then read anything into the rate of the varying ones.
  • How do you separate a feature that varies from one whose success rate has simply fallen?
    By the rate, measured over enough repeats. Variation is expected and its rate stays roughly stable; a fall shows as the same case meeting its criteria in a smaller fraction of repeats than before, on unchanged inputs. That comparison only exists if the per-repeat outcomes were kept, so today's figure has something to be read against.
  • The case is variable and also sits on a leaked fixture. How does that show up?
    As a diagnosis that never fully settles: pinning the generative step reduces the variation without removing it, and the remaining misses track the machine rather than the input. Treat them as two findings. Fix the test-side defect first, because it is the one that will otherwise corrupt every rate you measure afterwards.

saying these in an interview costs you the question

  • Calls every varying result a badly written test
  • Calls every varying result the feature's nature and moves on
  • Changes the criteria and the test code in the same edit
  • Never runs the case in isolation before blaming the feature
  • Ignores that the failing assertion was on a stored identifier
  • Stops at the first explanation that fits the symptom