In a Karate mock feature, how do you make one endpoint answer three seconds late, and why is that better than sleeping inside the scenario?
answer
- the mock serves one request at a time
- let the server wait, not the scenario
- another reply variable, set the same way
- the unit is milliseconds
basics
~20 sSet the variable: a step reading def responseDelay = 3000 in that scenario. The mock handler serialises requests, so sleeping inside a scenario stalls every other client, while responseDelay is applied by the server after the handler has already returned.
solid answer
~50 sYou add `* def responseDelay = 3000` to that scenario — milliseconds, read along with `response` and `responseStatus` once the last step has run, so its position in the scenario does not matter. The reason it beats a sleep is concurrency: a Karate mock handler processes one request at a time, which is what makes shared `Background` state safe, so a sleep inside a scenario holds the request slot and every other client waits behind it for the full three seconds. `responseDelay` is applied by the HTTP server **after** the handler has returned, on its event-loop scheduler, so only that one reply is held back. Note there is no `configure responseDelay` for a global default — it is only ever a per-scenario variable — and the delay is skipped on the 404 and 500 paths.
code
gherkin · 10 linesFeature: slow dependency mock
Scenario: pathMatches('/reports/{id}') && methodIs('get')
* def responseDelay = 3000
* def response = { id: '#(pathParams.id)', state: 'READY' }
Scenario: pathMatches('/reports')
# jitter around a client timeout instead of one fixed point
* def responseDelay = 200 + Math.random() * 400
* def response = []go deeper
Remember it is a variable in milliseconds set inside the scenario, like the body and the status, and that Karate applies it for you rather than you waiting in the steps.
Explain the mechanism: the handler serialises requests, so a sleep occupies the only slot while responseDelay is applied after the handler returns and blocks nothing.
Use it to reproduce a measured latency and to sit either side of a client timeout, and know the failure paths skip it — an instantly fast 'slow' route usually means a broken or unmatched scenario.
Decide what latency a shared mock should default to. Fast mocks make suites green and hide timeout bugs; realistic ones cost pipeline minutes. That trade belongs to whoever owns both the suite's runtime and its production incidents.
## One line, and it is a variable ```gherkin Scenario: pathMatches('/reports/{id}') * def response = { id: '#(pathParams.id)', state: 'READY' } * def responseDelay = 3000 ``` `responseDelay` is a number of **milliseconds**, and it is read alongside `response`, `responseStatus` and `responseHeaders` once the scenario's last step has finished. Like the others, where you put the line in the scenario does not matter — assembly happens afterwards — and a later step may overwrite it, which is how a scenario can decide the delay from the request it just parsed. ## Why it is not the same as sleeping This is the part that separates a correct answer from a plausible one, and it turns on how the mock handles concurrency. A Karate mock handler **serialises requests**: one request is processed at a time, so that the `Background` state every scenario shares cannot be torn by two requests at once. That guarantee is what makes a stateful mock — a counter, a store, a sequence of replies — safe to write in the first place, and it is also the reason a sleep is the wrong tool: - **Sleeping inside the scenario** happens *while the handler holds the request slot*. Every other client waits behind you for the whole three seconds. A "one slow endpoint" test quietly becomes a "the entire mock is slow" test, and a concurrency test built on it measures the mock, not the system. - **`responseDelay`** is applied by the HTTP server *after* the handler has returned, on the server's own event-loop scheduler. The handler is free again immediately; only that one reply is held back. Other requests are matched and answered normally while it elapses. So the delay you want to model — a slow dependency — is exactly what `responseDelay` produces, and exactly what a sleep does not. ## There is no `configure responseDelay` `configure cors` and `configure responseHeaders` are real server-wide settings you declare in the `Background`. It is a very natural leap to assume `configure responseDelay` exists for a global latency floor. **It does not.** `responseDelay` is only ever a per-scenario variable, and the two Karate lines get that wrong in opposite directions: - on **Karate 1.x** the key is accepted by the configuration parser and then never read by anything, so `* configure responseDelay = 2000` in a `Background` is a **silent no-op** — the server starts, every reply is instant, and nothing in the log mentions it; - on **Karate 2.x** it is not a known configuration key at all, so the step throws `unexpected 'configure' key: 'responseDelay'`, and a failing `Background` step is fatal: the mock server does not start. If you want latency on every route, set the variable in each scenario, or set it from a per-request hook, which runs before the reply is assembled and can therefore still change it. ## Where the delay does not apply `responseDelay` shares the success-path restriction with the configured headers and the CORS header: | outcome | delay applied? | |---|---| | matched scenario, all steps passed | yes | | matched scenario, a step failed (500) | no — the error comes back immediately | | no scenario matched (404) | no | That asymmetry is useful diagnostically. If the endpoint you slowed down suddenly answers instantly, the reply you are getting is probably not the one you think: a fast reply from a scenario that should be slow is a good sign that the scenario broke or never matched. ## Using it well 1. **Model a real number.** Delay is most valuable when it reproduces a measured p95 of the real dependency, not a round number chosen because it is memorable. 2. **Vary it when you are testing a timeout.** A constant delay tests one point; a computed one — `* def responseDelay = 200 + Math.random() * 400` — exercises the region around a client's timeout, where the interesting failures are. 3. **Keep it near the top of the scenario.** It is read at the end regardless, but a reader who scans the first two lines should be able to see that this route is deliberately slow. 4. **Delay is not a fault.** A slow reply and a broken connection fail a client in different ways, and `responseDelay` only models the first.
- Your slow endpoint suddenly starts answering instantly. What would you check first?Whether the reply is still coming from the scenario you think. `responseDelay` is applied only on the success path, so a failed step (500) or an unmatched request (404) comes back immediately with no delay at all. Check the status and the mock's log for a step failure or a "no scenarios matched" line before suspecting the delay itself.
- How would you give every route in a mock a baseline latency?Not with a configuration key — there is no `configure responseDelay`. Either set the variable in each scenario, or set it from a per-request hook, which runs before the reply is assembled and can therefore still write `responseDelay`. Keeping it in one hook also makes the whole mock's latency profile a single line to change.
- What does responseDelay not model?Anything that is not a slow but complete reply. A connection that is refused, reset mid-body or held open until the client's own timeout fires all fail a client differently from a reply that simply arrives late, and a delay reproduces none of them. Use it to exercise timeouts and retries, not to simulate a dead dependency.
A sleep is the clerk taking a nap at the counter — the whole queue waits. responseDelay is a timer on the outbox: your envelope is held back three seconds, and the clerk serves the next person immediately.
saying these in an interview costs you the question
- Reaches for a sleep or a busy loop inside the scenario instead.
- Assumes configure responseDelay gives every route a global delay.
- Thinks the delay is measured in seconds rather than milliseconds.
- Believes the mock handles other requests during a scenario-level sleep.
- Expects the delay to apply to an unmatched 404 or a failed-step 500.