How do you use a virtual service to test how a client behaves when its dependency is slow, throttled or failing?
answer
- Answer badly, on purpose, every run
- Delay just above the configured timeout
- A dead connection is not an error status
- Valid syntax can still be wrong shape
- Assert the client, not the stand-in
basics
~20 sConfigure the stand-in to answer badly: a delay just above the client's configured timeout, an error or throttling status, a dropped connection, or a malformed body. Then assert on the client's retries, error mapping and stored state.
solid answer
~50 sA virtual service can answer badly on purpose, and that is usually why it is there. The four families are **latency** (a delay tuned just above the client's configured timeout, so the timeout path runs without a long wait), **protocol-level faults** (connection refused, reset mid-response, a truncated body - paths a clean error status never reaches), **status responses** (unauthorized, throttled with a retry hint, server error), and **malformed payloads** (invalid syntax, valid syntax with a wrong field type, a missing required field). For anything involving recovery you need state: a scenario where the first two matched calls fail and the third succeeds proves the retry actually completes rather than merely being attempted. The assertions are about your side - how many calls were made, with what spacing, whether a retry reused the same idempotency value, what the user was shown, and above all whether a partial write was left behind.
code
pseudocode · 14 linesscenario("submission-retry")
.state("start")
.when(method = "POST", path = "/returns")
.respond(status = 503).then("attempt-2")
.state("attempt-2")
.when(method = "POST", path = "/returns")
.respond(status = 429, headers = { "Retry-After": "2" }).then("attempt-3")
.state("attempt-3")
.when(method = "POST", path = "/returns")
.respond(status = 202, body = { "receiptId": "R-84417" })
assert journal.callsTo("/returns").count == 3
assert journal.callsTo("/returns").allShare("Idempotency-Key")
assert storedReturn.status == "FILED" and storedReturn.receiptId == "R-84417"go deeper
Be ready to name the kinds of bad behaviour a stand-in can produce - a delay, an error status, a broken connection, a corrupt body - and to say why those paths are worth testing at all.
Explain the mechanics: sizing a delay against the client's configured timeout, why a dead connection is a different branch from an error status, and how a stateful scenario lets a call fail twice and then succeed.
Demonstrate that you assert on the client, not the stimulus: call count and spacing, backoff honouring a retry hint, a reused idempotency value, error mapping, and the persisted state left behind after an unknown outcome.
Own the question of how far this goes: which failure modes are worth encoding, how the team keeps injected faults from leaking across a shared stand-in, and when the organisation should be exercising real degradation in a controlled environment instead of a believed one.
## Why this is the most common reason to stand in a service The paths that break in production are the ones nobody exercises: the dependency was slow, it throttled you, it returned a body the parser had not seen. A real sandbox will not produce those on request. A virtual service will produce them on every run, deterministically, which turns resilience from an intention into a tested property. ## The four families ### 1. Latency Configure a delay on the matched response. The important detail is the *size* of the delay. A tax-filing wizard's submission client is configured with a 250 ms read timeout; the stand-in delays 420 ms. That reaches exactly the same code path as a thirty-second hang, and the 340-case regression pack stays fast. Injecting a long sleep instead is the classic novice move - it proves the same thing and costs the suite minutes. Delay variants worth knowing: a fixed delay, a delay applied to the first chunk only (slow to respond) versus spread across the body (slow to transfer, which some clients time out differently), and a random delay inside a band, which is useful for soak runs but a poor choice inside a deterministic pack. ### 2. Protocol-level faults A clean error status is a *successful* exchange carrying bad news, and it exercises a different branch from a connection that never opened or died mid-flight. Worth injecting: refusing the connection, accepting then closing without a response, closing after headers but before the body ends, and returning garbage bytes where a response was expected. The truncated-body case is the one that catches the wizard-shaped defect, because a parser handed half a document may return a partially populated result rather than raising - which is how a silent data corruption starts. ### 3. Status responses Unauthorized and forbidden (does the client refresh a credential, or retry forever?), not found, conflict, throttled with a retry hint, and the server-error family. Throttling deserves its own case: return the throttling status with a retry hint of two seconds and assert that the client waits rather than hammering, and that its retry budget is finite. A client that treats throttling as a generic failure and retries immediately turns a provider's brownout into an outage. ### 4. Malformed payloads Three distinct rungs, and candidates usually name only the first: - **Syntactically invalid.** Truncated or corrupt bytes. The deserializer should raise. - **Syntactically valid, semantically wrong.** A number where a string was expected, an unknown enumeration value, a null in a required position. This is where lenient parsers quietly produce a default value. - **Valid and plausible, but wrong for the request.** A receipt for a different submission identifier. Only an assertion that correlates request to response will catch it. ## Stateful scenarios Single-response rules cannot express recovery. A stand-in that supports scenarios lets a rule advance a state on each matched call: state *start* returns a server error and moves to *attempt-2*, which returns a throttling status and moves to *attempt-3*, which returns the receipt. Now the case proves the whole retry sequence - that it recovers, that it stops in time, and that it does not file the return twice on the way. The same mechanism expresses a dependency that is healthy, then degrades, then recovers, which is how you test a circuit-breaking client without waiting on real weather. ## What you assert The mistake is to assert on the stand-in. The stand-in is the stimulus; the client is the subject. Assert: - **Call count and spacing** from the request journal - three attempts, not eleven, with increasing gaps. - **Idempotency.** The retry carried the same idempotency value, so the gateway can deduplicate. - **Error mapping.** A throttle surfaced as *try again shortly*, not as a stack trace or a validation message. - **State after failure.** This is the one that matters most in the tax-filing example: after the timeout, is the return marked submitted locally? A wizard that records a submission it never confirmed is the silent data corruption, and only an assertion on persisted state after an injected fault will find it. ## Practical cautions Keep injected-fault cases in their own group; mixing them with happy-path rules on one shared stand-in leads to a fault leaking into unrelated cases. Make the injected behaviour explicit in the case name, so a failure reads as *timeout path* rather than as a mystery. And remember the limit of the technique: you are testing your client's response to a *believed* failure mode. That the real dependency actually fails that way is a separate claim, and it decays like any other.
- Why inject a delay just above the configured timeout instead of a long fixed pause?Because the client cannot tell the difference - once the deadline passes, the same branch runs - and the pack keeps its speed. With a 250 ms timeout, a 420 ms delay proves the timeout path as well as thirty seconds would, and across a 340-case pack that choice is the difference between minutes and an hour. Leave the long hangs for a deliberate soak run.
- Which failure modes cannot be expressed with an error status code at all?Anything below the response: a refused connection, a connection accepted then closed with no response, a body truncated mid-stream, a handshake failure, or bytes that are not a valid message at all. These reach different branches in the client - transport error handling rather than status handling - and the truncated body in particular can leave a lenient parser returning a half-populated result instead of raising.
- After injecting a timeout on a submission call, what is the single most valuable thing to assert?The persisted state on your own side. A timeout means the outcome is unknown, so the record must not claim the submission succeeded; it should sit in an explicitly unconfirmed state that a reconciliation path can resolve. Asserting only that an error was surfaced misses the corruption, because the damaging bug is a local record marked complete for a call that may never have landed.
saying these in an interview costs you the question
- Injects a thirty-second sleep to exercise a timeout
- Treats a dropped connection and a server error as the same path
- Retries a throttling response immediately and calls it resilience
- Only tests invalid syntax, never a wrong field type
- Asserts the stand-in returned an error, not what the client did
- Never checks persisted state after an injected failure