skip to content

The same build of an AI-powered feature answers one user correctly and another wrongly. What differences do you rule out first?

level: middleimportance: should knowfreq 45%

answer

  1. same build is not the same request
  2. compare the two reconstructed requests
  3. accounts assemble different material
  4. rollouts make two products, one build
  5. variation is what remains, not first guess

basics

~20 s

Rule out deterministic differences before blaming any layer: the literal inputs each user sent, the material each account's request assembled, the variant and settings served, and the client each used. Only then consider run-to-run variation.

solid answer

~50 s

Same build does not mean same request. Work through the differences that are cheap and deterministic to check: - **The literal inputs.** Two users who describe the same question rarely send the same text. Compare what was actually sent, not the summaries in the two reports. - **The material each request assembled.** It usually depends on the account: permissions, organisation, history, entitlements. One user's supporting document may be missing, out of date, or invisible to them. - **The variant and settings served.** Gradual rollouts, per-region configuration and per-cohort model versions all produce two different products behind one build number. - **The client and session.** A stale client, an older cached answer, or locale-dependent formatting can differ while the server is identical. Only when all four match does the difference point at the generative step's own run-to-run variation, and that is confirmed by repeating one identical call, not by comparing two users.

code

pseudocode · 15 lines
pseudocode
a = reconstructCall(complaintFromUserA)   // wrong answer
b = reconstructCall(complaintFromUserB)   // correct answer

diffs = []
if a.inputText        != b.inputText:        diffs.add("input text")
if a.suppliedMaterial != b.suppliedMaterial: diffs.add("assembled material")
if a.variant          != b.variant:          diffs.add("variant assignment")
if a.settings         != b.settings:         diffs.add("settings in force")
if a.clientVersion    != b.clientVersion:    diffs.add("client or cache")

if diffs not empty:
    investigate(diffs.first)              // deterministic, has an owner
else:
    outcomes = repeat(a, 20)              // identical requests, both outcomes?
    report("one feature-wide failure rate: " + failureRate(outcomes))

go deeper

for a junior

Be ready to list what can differ between two users on identical code: the words they typed, the data their account can reach, the variant they were served, and the client they read the answer on. Same build is not the same request.

for a middle

Explain how you separate those cheaply: reconstruct both calls, send one user's exact text on the other's account, compare the settings actually in force, and only then repeat an identical call to test for run-to-run variation.

for a senior

Show the judgment about which error is more expensive. Attributing a real defect to variation hides it forever; chasing a per-user cause that does not exist burns days. Both are avoided by reconstructing both requests fully before forming a theory.

for a principal

Own what makes this answerable at all. If requests are not reconstructable per user, no attribution is possible, and gradual rollouts guarantee that one build number covers several different products. Decide what has to be recorded so two complaints can be compared at all.

## Works for me means a difference you have not found yet Two users, one build number, two different outcomes. The instinct is to treat this as evidence about the generative model, because a component that can answer differently each time is an obvious explanation. It is also the last one that should be reached for, because four cheaper and entirely deterministic differences sit in front of it, and each of them is checkable in minutes. The framing that helps is this: a build number describes the code that was deployed, and almost nothing else that decides what a user gets. What actually reaches the generative step is a request assembled at run time from the user's text, their account's data, their assigned variant and their session state. Two people on identical code can be sending materially different requests, and usually are. ## The differences worth ruling out | Difference | How to rule it out | What it would mean | | --- | --- | --- | | The literal input text | Compare the two recorded inputs character by character, not the two complaint summaries | Different questions were asked, and no layer is at fault yet | | The material each request assembled | Read the material as sent for both, and check what each account is permitted and able to reach | The assembly of material owns it: missing, stale, or restricted for one account | | The variant and settings served | Check the rollout assignment, region configuration and model version behind each request | Configuration owns it, and an untested combination is reaching users | | The client and session | Compare client versions, cached answers and locale-dependent formatting | The failure is downstream of the answer, in ordinary product code | | Run-to-run variation of the generative step | Only after the four above match: repeat one identical call many times | There is no user difference at all; you are looking at a rate | The order matters because each row is cheaper than the one below it and can end the investigation on its own. It also matters because rows one to four produce *fixable* findings with named owners, while row five produces a rate that has to be managed rather than repaired. ## Ruling them out cheaply 1. **Get both recorded requests side by side.** Input text, assembled material, settings, returned text, shown answer. Comparing two reconstructions is far faster than reasoning about two reports. 2. **Normalise the inputs and retry.** Send the failing user's exact text on the working user's account, and the reverse. If the failure follows the text, the account is not the variable; if it follows the account, the material or the variant is. 3. **Check what each account can reach.** Permissions and organisation boundaries silently change what material is assembled, and this is the difference most often missed because the code path is identical for both. 4. **Compare the settings actually in force**, not the ones in the repository. Gradual rollouts exist precisely to make two users different, and they do that job perfectly. 5. **Only then repeat one identical call.** If the same account, text, material and settings produce both outcomes across repeats, the difference between the two users was never real; it was two samples from one distribution. ## When there is genuinely no difference Sometimes step five is where you land, and the honest statement is that the feature fails at some rate for everyone and two users happened to sit either side of it. That changes the shape of the work completely. There is no per-user bug to fix; there is a rate to measure, a containment to design, and a decision to make about a feature that is only usually right. It also means the two complaints are one complaint, and treating them as two independent defects will produce two independent, contradictory investigations. ## What the answer changes The practical value of this discipline is that it protects against the two errors that cost the most. The first is attributing a real, fixable defect to run-to-run variation, which closes the ticket, leaves the defect in place, and teaches the team that this feature is simply unpredictable. The second is the mirror image: chasing a per-user cause through days of investigation when the two users differed in nothing at all, and inventing an explanation to fit. Both are avoided by the same cheap habit of reconstructing both requests fully before forming any theory about which layer is responsible.

  • Which of these differences is missed most often, and why?
    What each account is permitted and able to reach. The code path is identical for both users, so nothing in the request handling looks different, yet the material assembled for one of them is missing a document the other can see. It is invisible in logs that record the code path rather than the material actually sent.
  • Both users sent identical text on identical settings and one still failed. What is the next move?
    Stop comparing the two users and start measuring one request. Repeat the identical call many times and record how often the answer is unacceptable. If both outcomes appear, the two complaints are one complaint about a feature-wide rate, and the work moves from finding a per-user cause to measuring and containing that rate.

Two people receive different bills from the same billing system: before suspecting the arithmetic, you check that they were on the same tariff, same period and same meter.

saying these in an interview costs you the question

  • Assumes the same build means the same request
  • Blames run-to-run variation before comparing the two requests
  • Compares the two complaint reports instead of the recorded calls
  • Forgets that accounts assemble different supporting material
  • Treats two samples of one failure rate as two separate defects