What version assumptions does an end-to-end run make when components deploy independently?
answer
- A result is bound to more than code
- Which combination was actually deployed
- Production rarely runs latest of everything
- A deploy landing mid-run
- Read build identifiers at both ends
basics
~20 sAn end-to-end run is evidence about one combination of deployed component versions at one moment. Where components release independently that combination may never exist in production, and can even change mid-run, so record every version with the result.
solid answer
~40 sA green journey means *this set of builds worked together, in this environment, at this time*. Independent deployment breaks each assumption in turn. The combination tested is whatever happened to be deployed - usually the newest of everything, which production may never run, since production carries a mix of newer and older components. A deploy landing mid-run splits the results across two combinations nobody can interpret. And a rolling release can serve two versions of one component at once, so a journey may cross both. So **stamp every build identifier into the report**, **pin the combination for the run window**, and accept that this level cannot answer whether one component is safe to release alone - that needs per-pair compatibility evidence, a different artefact entirely.
code
pseudocode · 10 linesbefore = readBuildIds(["seat-map", "inventory", "pricing", "booking"])
report.header(environment, before, startedAt = now())
runAllJourneys()
after = readBuildIds(["seat-map", "inventory", "pricing", "booking"])
if after != before:
report.invalidate("deploy landed mid-run: " + diff(before, after))
else:
report.publishVerdict()go deeper
Be ready to say that a passing journey is a statement about specific deployed builds, not about a source branch, and that the report should record which builds those were.
Explain why a run against the newest of everything can miss defects: the mixed combinations that production actually serves were never exercised, and a rolling release can even put two versions of one component in the same journey.
Show that you make the run interpretable - build identifiers read before and after, a deploy freeze or a pinned deployment for the run window, and a report that refuses a verdict when the set changed underneath it.
Own the boundary of the level: it cannot answer whether one component is independently releasable against its neighbours' real versions, and be ready to say what artefact you would fund for that question instead of multiplying journeys.
### What a green run is actually evidence of An end-to-end result is bound to four things: the set of component versions deployed, the environment they ran in, the data present, and the moment in time. Interview answers usually remember the first and forget that it is a *set* - and a set that no one deliberately chose. When a single artefact is deployed as one unit, the version question is trivial: one build, one result. As soon as components release independently - an airline seat-map front end, an inventory component, a pricing component and a booking component, each with its own release cadence - the run tests whichever combination happened to be live. Three distinct assumptions hide in there. ### Assumption 1: the combination under test is a combination that will exist A shared pre-production environment usually drifts toward latest-of-everything, because every team deploys there first. Production rarely looks like that: one component is two releases behind because its rollout was paused, another is a release ahead because it shipped a fix on Tuesday. A journey that is green against the newest of everything can still break in production on a combination nobody exercised. The inverse also happens - a journey fails in the shared environment against a combination that will never ship, and a team spends a day on a defect that could not reach a user. ### Assumption 2: the combination is stable for the duration of the run If the seat-map front end is redeployed eleven minutes into a nineteen-minute suite, the first half of the results describe one system and the second half another. The report shows a single verdict for two different systems. Worse, the failure looks exactly like nondeterminism, so it gets attributed to test unreliability rather than to a mid-run deploy - and the real signal, that two versions disagree, is thrown away. Mitigations are procedural as much as technical: run against a deployment that is pinned for the window, gate deploys to the target environment while a run is in flight, or run against an environment stood up for that run from a named set of versions. Each has a cost - contention, provisioning time, complexity - and the choice depends on how expensive an uninterpretable run is to the team. ### Assumption 3: one component means one version A rolling or staged release has two versions of the same component serving at once, behind whatever routes traffic. A journey with several requests may cross both, so a step recorded against the new build and a later step against the old one. If the two disagree about a payload shape or a session representation, the journey fails in a way that reproduces only under that routing. This is real behaviour that production also has, so it is not purely a testing artefact - but it must be recognised, or it reads as an inexplicable one-off. ### What to record, always The cheapest and highest-value practice is to stamp the report: **every component's build identifier, the environment name, and the run start and end timestamps**, printed in the header of the result. It costs a small amount of plumbing - each component exposes its build identifier, the suite reads them before and after the run - and it answers the first question of every investigation. Reading them again at the end also detects a mid-run deploy automatically: if the identifiers differ from the ones read at the start, the run is uninterpretable and should say so rather than report a verdict. ### The limit of the level, and what covers it instead The question this level structurally cannot answer is: *can this one component be released on its own, against the versions its neighbours are actually running?* A journey only ever exercises the combination in front of it, so answering that would need a run per candidate combination, and the number of combinations grows with the number of components. Recognising that the combinatorial question needs a different artefact - pairwise compatibility evidence recorded per consumer and provider, checked at build time - is the senior move; the details of that artefact belong to a different discipline and should be handed to it rather than reinvented here. ### How to answer this in an interview Say what the result is evidence of, in one sentence, and then name the three assumptions and their mitigations. The differentiating detail is the mid-run deploy: most candidates never consider that a suite can span two systems, and the fix - read the build identifiers at both ends of the run and refuse to report a verdict when they differ - is concrete, cheap and clearly the answer of someone who has been burned by it.
- A journey passes in the shared pre-production environment and fails in production. What do you check first?Whether the two ran the same combination of component versions. Pre-production drifts toward latest-of-everything while production carries a mix, so a defect that only appears against an older neighbour is invisible where you tested. Compare the recorded build identifiers from the green run against what production is actually serving; if they differ, the question becomes which pair is incompatible, not why the test was wrong.
- Is running against latest-of-everything a defensible policy?It is defensible as the default because it is cheap and catches integration breakage early, but it should be stated as a policy rather than assumed. Its blind spot is the mixed combinations production actually runs, so a team relying on it needs some other evidence for backward compatibility across a release boundary. Whether it is enough depends on how far components are allowed to drift apart in production.
Testing one combination of independently released components is like certifying a chord: the notes sounded good together this once, which says nothing about the chord the orchestra will actually play tonight.
saying these in an interview costs you the question
- Treating a green run as a property of the code alone
- Never recording which versions the run exercised
- Assuming pre-production and production run the same combination
- Blaming a mid-run deploy on test unreliability
- Believing one journey certifies a single component as releasable
- Ignoring that a rolling release serves two versions at once