How do you correlate a failed automated case with the system's own record of the same request?
answer
- Two records of the same event need a key
- Timestamps and guesswork do not scale
- The case mints something the system carries
- Attach it to every outbound call, not the last
- Print it beside the failure evidence
basics
~20 sMint an identifier per case, attach it to every outbound call in a field the system propagates through its own layers and writes on its log lines, and print it with the failure. The report then points straight at the system-side records.
solid answer
~50 sMake the case emit an identifier the system will carry. At the start of each case the harness mints a value that is unique to that case in that run, and attaches it to every outbound call the case makes, in whatever field the system already propagates through its own layers. The system writes that value on every log line it produces while handling those calls and carries it across service boundaries. On failure the harness prints it alongside the evidence, so whoever reads the report can pull the exact system-side records for that one execution instead of guessing from timestamps. Timestamps alone do not work: parallel workers overlap, clocks drift, and a busy target logs thousands of similar requests a second. Retrofitting this means agreeing on one propagated field, not instrumenting each case by hand.
code
pseudocode · 15 lines# once, in the harness - every case inherits it
before_case(case):
case.correlationId = join("case", case.id, run.id, randomSuffix())
outboundClient.defaultFields["correlation-id"] = case.correlationId
after_case(case, outcome):
if outcome.passed:
return
bundle.put("correlationId", case.correlationId)
bundle.put("runId", run.id)
report.line("system records for this failure: correlation-id=" + case.correlationId)
# optional: the system's own retention is shorter than triage takes
if systemLogs.reachable():
bundle.put("systemLines", systemLogs.fetchBy(case.correlationId, window = 5min))go deeper
Know the idea before the plumbing: a failed case and the system's own logs are two accounts of the same event, and something shared has to link them. Be able to say why searching by the time of failure is unreliable.
Explain the mechanics — where the value is minted, which outbound calls carry it, how the system copies it onto every line it writes, and why the harness prints it with the failure rather than leaving it in a variable nobody sees.
An interviewer at this level expects the retrofit story: you do not own the system's logging, so you argue for one already-propagated field, handle the case where a value arrives from the caller, and prove the trail survives every boundary the work crosses.
Own it as a platform property rather than a testing trick. Decide who is accountable for propagation, what a service must do to count as diagnosable, and how you fund that work when the benefit lands on somebody else's on-call rota.
## Two records of the same event When an automated case fails against a running system, two independent accounts of the same work exist. The harness has one: the steps it performed, the responses it saw, the observation that broke. The system has the other: its own log lines, its own timing, its own internal errors, written by code the harness never touches. The failure that matters usually lives in the second account — a validation the case never sees, a dependency that returned an error the entry layer swallowed, a slow query. Joining those two accounts is the difference between "the total was wrong" and "the discount service timed out and the entry layer defaulted to zero". The join needs a key. Providing one is the whole job. ## Why time-based matching fails Matching by timestamp is the default when nobody has designed anything better, and it degrades badly: - **Concurrency.** A parallel run has many workers issuing similar requests in the same instants. Time narrows the search to hundreds of candidates, not one. - **Clock skew.** The machine running the suite and the machines running the system do not agree to the millisecond, and often not to the second. - **Volume.** A busy target logs thousands of lines a second, most of them from traffic that has nothing to do with the run. - **Asynchrony.** Work that continues after the call returns is logged at a time unrelated to the case's timestamps. The failure mode is not that time-based matching never works — it is that it works on a quiet system and stops working on the busy, parallel, production-like one where the interesting failures actually happen. ## What the identifier has to satisfy 1. **Minted by the case, not by the system.** If the system generates it, the harness has to read it back out of a response, which fails the moment a call errors before the value is returned. 2. **Unique to one execution of one case.** Reused across a run it identifies nothing; reused across runs it makes yesterday's failure look like today's. 3. **Carried on every outbound interaction the case makes**, not only the one the failing observation checked — because the cause is frequently several steps earlier. 4. **Propagated by the system through its own layers**, including across service boundaries and across any handoff to work that continues after the response is sent. 5. **Printed with the failure**, not merely held in a variable. An identifier the reader never sees links nothing. ## Which identifier answers which question | Identifier | Scope | What it answers | |---|---|---| | Run value | The whole suite execution | Which run produced this evidence | | Case value | One case, stable across runs | Which behaviour was being checked | | Correlation value | One case in one run | Which system-side records belong to this failure | | Per-interaction value | One call | Which single interaction inside the case | Keep the run value and the correlation value both, and derive the second from the first so they can never disagree. The run value groups a run's evidence; the correlation value is what actually finds the failure. ## Retrofitting propagation Most teams do not get to design this from scratch, and the retrofit is a system change rather than a harness change. The move is to find the one field that already crosses every boundary — most systems have some inbound value the entry layer generates for its own logging — and make three amendments: - Accept a caller-supplied value when one is present, generate one when it is not. - Write it on every log line produced while handling that work, not just the first. - Copy it onto anything enqueued or handed off, so the trail does not stop at the boundary where the response is returned. Resist inventing a parallel channel that only the tests use. A field only the suite populates rots, because nothing in production depends on it. A field the system already relies on for its own diagnosis is maintained by people who will notice when it breaks. ## Diagnostic, not a verdict One boundary is worth stating plainly. The identifier exists so a human can find the system's account of a failure after the fact. It is not what decides whether the case passed — the case's own observations do that. Keeping those separate matters: a harness that starts reading system-side records to decide outcomes has quietly moved the definition of correct out of the case and into whatever the system happened to log. ## Where it goes wrong - The value is minted but never printed with the failure, so it helps nobody. - It is attached only to the final call, so failures caused three steps earlier are invisible. - The system receives it and does not log it, which looks identical to the trail not existing. - The system's records are retained for less time than it takes anyone to read the report. Where that is true, have the harness pull the small window of records for its own identifier at failure time and fold them into the evidence, so the report stands on its own.
- The system propagates no identifier you control. What is your first move?Make an existing field carry it rather than adding a parallel channel only the suite uses. Most systems already generate some inbound value for their own logging; the change is to accept a caller-supplied value when present, generate one otherwise, write it on every line produced while handling the work, and copy it onto anything handed off to continue later.
- Why is a per-case identifier better than a per-run one for diagnosis?A run value narrows the system-side records to the whole suite, which on a parallel run is thousands of interleaved requests. A per-case value isolates one story. Keep both — the run value groups the evidence, the case value finds the failure — and derive one from the other so they can never disagree.
- What do you do when the system's records are kept for less time than it takes to read the failure?Copy what you need at failure time instead of relying on the source. The harness can pull the small window of records matching its own identifier while they still exist and fold them into the evidence it writes, so the report stays self-contained even after the original is rotated away.
It is the tracking number on a parcel: everyone who handles it writes down the same number, so one lookup reconstructs the whole journey instead of searching for a box that left on a Tuesday.
saying these in an interview costs you the question
- Matches the failure to system logs by timestamp alone
- Mints an identifier but never prints it with the failure
- Adds the identifier by hand in each case that needs it
- Assumes the system will log a value it never received
- Attaches the value only to the last call the case makes
- Uses one value for the whole run and none per case