How do you detect that a virtual service has drifted from the real dependency it stands in for?
answer
- The stand-in froze; the provider did not
- False failures are loud, false passes silent
- Re-record on a cadence and diff
- Validate canned responses against the published schema
- Every stand-in needs an owner and an expiry
basics
~20 sRe-record real traffic on a schedule and diff it against the canned responses, validate every canned response against the provider's published schema in the pipeline, and keep a small tier of checks running against the real service so the belief is re-earned.
solid answer
~50 sDrift is a stand-in that no longer resembles the service it imitates, and it is dangerous because it fails *green*. Three mechanisms catch it. First, scheduled re-recording: route a representative slice of the suite through a recording proxy against the real sandbox and diff the captures against the current canned responses, with volatile fields excluded so the diff means something. Second, validation in the pipeline: check every canned response against the provider's published schema, so a field that changed type fails a build rather than a filing. Third, a deliberately small live tier - a handful of checks that call the real service on a cadence and assert only on shape. Around those, put governance: an owner per stand-in, the provider's specification revision recorded next to the rules, and a non-zero unmatched-request count treated as a failure. Also watch production for parse errors and unknown fields, which are drift observed too late.
go deeper
Be ready to say what drift is in one sentence - the stand-in froze a belief and the real service moved on - and why that shows up as tests passing rather than failing.
Explain at least two detection mechanisms and what each cannot catch: schema validation misses meaning changes, re-recording needs volatile fields scrubbed before the diff is readable.
Show the asymmetry - false failures are loud and cheap, false passes are silent and expensive - and describe a concrete routine you have run: cadence, what gets re-recorded, how a diff is triaged, and how production parse errors feed back into the rules.
Own it as policy: which dependencies may be virtualized, who owns each stand-in, what expiry and evidence rules apply to a rule edit, and how you would argue the organisation's case for a provider-maintained sandbox rather than dozens of independently rotting imitations.
## What drift is, and why it is the leaf's whole point A virtual service encodes a belief about another team's service on the day someone wrote or recorded it. The other team never agreed to freeze. Drift is the gap that opens afterwards, and its danger is asymmetric: - **Your stand-in is stricter than reality.** You get false failures. Expensive in attention, but loud, and someone fixes it the same day. - **Your stand-in is looser or simply wrong.** You get false *passes*. A 340-case regression pack goes green while the behaviour it claims to protect is already broken in production. That is confidence you did not earn, and nothing in the pack will ever tell you. The tax-filing example: the gateway migrated a monetary field from an integer in minor units to a decimal string. The stand-in kept serving integers. The wizard's parser, being tolerant, coerced the new decimal string on the real gateway and dropped the fractional part - a silent data corruption on live returns, discovered by a reconciliation report weeks later, with a fully green pipeline throughout. ## Where drift comes from 1. **The provider changed.** A new field, a widened enumeration, a changed type, a status that now means something else. 2. **The stand-in was edited to make a test pass.** Someone loosened a rule or hand-wrote a response that the real service has never produced. This is the most common source and the hardest to see in review, because the diff looks like a small configuration change. 3. **The recording aged.** Captures made once, eighteen months ago, by someone who has left. 4. **Your own client changed** and its requests no longer match rules that were once precise - which shows up as a rising unmatched count rather than as a failure. ## Detection, from cheapest to strongest ### Schema validation in the pipeline If the provider publishes a specification, validate every canned response against it as a build step, and validate the *recorded requests* against it too. This is cheap, runs on every change, and catches type changes, removed fields and invalid enumeration values the moment the specification is refreshed. It cannot catch semantic drift - a field whose meaning changed while its type did not - so it is a floor, not a ceiling. ### Scheduled re-recording and diff The strongest routine signal. On a cadence, run a representative slice of cases through a recording proxy against the real sandbox, capture the exchanges, and diff them against what the stand-in currently serves. The engineering work is in making the diff meaningful: scrub or ignore timestamps, identifiers, trace headers and ordering that does not matter, or the output is unreadable and gets muted. When a run against 1,412 recorded interactions produced 14 differences, 11 were volatile fields and 3 were real shape changes - and that ratio is typical, which is why triage rules matter more than the capture itself. ### A small live tier Keep a handful of checks that call the real service on a schedule and assert only on the shape of what comes back: status, presence and type of the fields you consume, one representative error. They are slow, they need real credentials, and they will occasionally fail for reasons that are not your fault - so keep them few, keep them out of the pull-request path, and treat a failure as a signal to refresh the stand-in rather than as a broken build. ### Signals you already have - **Unmatched-request count.** Rising means your client changed or the rules did not keep up. - **Production telemetry.** Deserialization errors, unknown-field counts, and unexpected-status counters from real traffic are drift already observed, just late. Feeding those back into new stand-in rules closes the loop. - **Provider communication.** Changelogs, deprecation notices and a named contact. Unglamorous and often the earliest warning available. ## Governance that makes the mechanisms stick Detection decays too, unless someone owns it. Practical habits: name an owner for each stand-in; record the provider's specification revision or sandbox version alongside the rules, and alert when it moves; stamp recordings with a capture date and treat anything past a chosen age as expired; and require that a rule edited purely to make a failing case pass carries evidence - a capture, or a line from the specification - that the real service behaves that way. ## The judgement to state out loud You are choosing how much of your confidence to borrow. A stand-in buys speed, determinism and unreachable states, and it pays for them with a claim about someone else's system that is true only until it is not. The mature position is neither 'virtualize everything' nor 'only test against the real thing' - it is that every stand-in has an expiry, a detection mechanism, and an owner, and that a green suite built on a stand-in nobody has re-validated in a year should be described honestly as an untested system.
- A scheduled re-recording diff reports 14 differences. How do you triage them?Separate volatile from structural. Timestamps, generated identifiers, trace headers and insignificant ordering are noise and should be added to the scrub list so they never appear again. What remains - a new field, a changed type, a widened enumeration, a different status for the same request - is real drift, and each item becomes either a stand-in update plus a client change, or a deliberate decision to ignore it, recorded.
- Someone edited a stand-in rule to make a failing case pass. How do you stop that from becoming drift?Require evidence in review: a capture from the real service or a citation from its published specification showing the response is real. Keep hand-invented responses in a separate, clearly labelled group so nobody mistakes them for recorded truth. And watch for the tell - a rule change and a source change landing together in one commit, where the rule change is what made the build go green.
- Why is schema validation of the canned responses not sufficient on its own?Because it only sees structure. A field that keeps its type while its meaning changes - a status value reused for a different outcome, an amount that switched currency basis, a list that is now ordered - passes validation and still corrupts behaviour. Structural checks are a cheap floor that runs on every change; catching semantic drift needs periodic contact with the real service.
- How would you decide how many checks to keep running against the real dependency?Few enough that they stay fast, credentialed and unflaky; many enough to cover each response shape you actually consume plus one representative error. Coverage of shapes matters far more than coverage of cases: one call per distinct response document is usually a better spend than twenty variations on the same happy path. Run them on a cadence, off the pull-request path.
It is a map of a city that stopped being updated. Every journey you plan on it looks fine, and you only discover the street was closed when you are standing in front of the barrier - by which time you have already told everyone the route was safe.
saying these in an interview costs you the question
- Assumes a green suite proves the stand-in is still accurate
- Never re-records after the initial capture
- Edits a rule to make a case pass, with no evidence
- Keeps no live checks against the real dependency at all
- Treats every diff in a re-recording as a failure, then mutes them all
- Cannot name an owner or an expiry for a stand-in