In a Gatling simulation, how would you decide between retrying a failing step, dropping the affected virtual user, and stopping the whole run?
answer
- Three widths: block, user, run
- Retry, drop, or abort
- exitHereIfFailed ends only that user
- stopLoadGenerator ends it for everyone
- An abort is not a verdict
basics
~20 sMatch the block to the blast radius: tryMax retries one chain, exitHereIfFailed ends that virtual user's scenario, and stopLoadGenerator ends the run for everyone. Retry transient faults, drop users whose remaining work is meaningless, stop only when the run is void.
solid answer
~40 sGatling gives you three widths of guard, and the decision is which width the failure deserves. `tryMax(n)` and `exitBlockOnFail` act on one wrapped chain for one virtual user. `exitHereIfFailed` ends that user's scenario when its session is currently failed, leaving the rest of the population running. `stopLoadGenerator(message)` ends the whole run gracefully for everyone, and `crashLoadGenerator` is the failing flavour. Retry only what is transient and idempotent; drop a user when its remaining requests would be meaningless, such as one that never got a token; stop the run only when the measurement itself is invalid, like the wrong environment or a missing seed dataset. Each choice costs something: retries inflate counts, drops quietly change the load you applied, and a stop means the profile you declared is not the profile you ran.
code
java · 12 linesScenarioBuilder scn = scenario("checkout")
.stopLoadGeneratorIf(
"catalog was not seeded, aborting",
session -> session.getString("catalogSize") == null
)
.tryMax(2).on(
http("refresh token").post("/oauth/token")
)
.exitHereIfFailed()
.exec(
http("checkout").post("/checkout")
);go deeper
Be ready to name the three blocks and say which one affects a chain, which a user, and which the whole run.
Explain that exitHereIfFailed fires on that user's own failed session status, and that stopLoadGenerator hands a stop command to the run controller rather than ending one user.
Show the operational reasoning: what each guard costs the measurement, and where in the chain you would place it so it fires on the failure you meant.
Own the policy across a suite: what may be retried at all, when a run is worth aborting, and how a run whose guards fired is reported so nobody reads it as a clean result.
## Three blast radii Gatling's failure blocks are not interchangeable, and the useful way to hold them apart is by **how much of the run each one affects**. | block | what it ends | who it affects | |---|---|---| | `tryMax(n)` | one attempt at the wrapped chain, then retries it | this virtual user | | `exitBlockOnFail()` | the rest of the wrapped chain, no retry | this virtual user | | `exitHereIfFailed()` | this user's scenario, from that point on | this virtual user | | `exitHere` / `exitHereIf(condition)` | the same, unconditionally or on a condition you write | this virtual user | | `stopLoadGenerator(message)` | the whole run, gracefully | **everyone** | | `crashLoadGenerator(message)` | the whole run, as a failure | **everyone** | `stopLoadGeneratorIf` and `crashLoadGeneratorIf` are the conditional forms, taking the message first and the condition second. ## What each one actually does - **`tryMax`** re-runs the wrapped chain and resets the user's status between attempts, so a recovered user continues normally. Every attempt is still recorded. - **`exitHereIfFailed`** is `exitHereIf` with one fixed condition: *this user's session is currently failed*. The user leaves the scenario at that point; any open group is closed on the way out, so the results stay well-formed. - **`stopLoadGenerator`** sends a stop command to the run's controller and does not hand the session on. The controller cancels the run's own timers and stops every population; the message you pass is logged with the reason. The `crash` flavour is the same mechanism with a failing outcome instead of a clean one. ## How to decide 1. **Is the failure transient and the step idempotent?** Then retry it, with the smallest possible block around it. A token refresh, a read that occasionally times out — yes. A payment, an order, anything that mutates — no. 2. **Would this user's remaining requests still mean anything?** A user with no token will generate a run of authorization failures that tell you nothing about the system's performance. Drop it with `exitHereIfFailed` and keep the noise out of your results. 3. **Is the run itself invalid?** Wrong environment, an expired shared credential, a seed dataset that is not there. Only then `stopLoadGenerator` — and prefer the graceful flavour unless you want the run to end marked as a failure. 4. **Is none of the above true?** Then let it fail. A failure that is recorded and left alone is data; a failure you papered over is not. ## What each choice costs the measurement - **Retrying** raises both your request count and your error count above what the injection profile describes. The run is still readable, but only if you say which errors were recovered. - **Dropping users** silently changes the load you applied *after* that point. The population you declared is not the population that reached the later steps, and nothing in the report announces it. - **Stopping the run** is the most expensive: the profile you declared is not the profile you ran, and the run is not comparable with the ones before it. Treat it as an abort, not as a verdict — Gatling has its own mechanism for deciding whether a completed run passed. ## Where the guards go in the chain Placement is half the decision. A guard only sees what has already happened to that user's session, so: - Put `tryMax` **around the smallest chain that can fail together** — the refresh call, not the refresh call plus everything it enables. - Put `exitHereIfFailed` **immediately after** the block whose failure should end the user. Further down the chain it will also catch failures from steps in between, which is rarely what you meant. - Put a `stopLoadGeneratorIf` **before the work**, typically as an early sanity check on the environment, so it fires once at the start rather than at an arbitrary point in a run you have already half-paid for. ## Anti-patterns worth naming - **Using `stopLoadGenerator` as a quality gate.** It ends the measurement instead of judging it, and it does so at whatever moment the first user tripped the condition. - **Blanket `tryMax` around everything.** It converts a visible failure rate into an invisible one plus extra traffic. - **Putting `exitHereIfFailed` inside a loop body expecting it to break the loop.** It does not break the loop; it ends the user's scenario outright. - **Retrying with no bound on the cause.** If the backend is failing every call, `tryMax(5)` turns one failed user into five failed requests and multiplies the load you are applying to something already unhealthy. - **Choosing the crash flavour reflexively.** Reserve it for conditions where a run that carried on would be worse than no run.
- Why is stopLoadGenerator a poor way to make a run fail on too many errors?It aborts the measurement instead of judging it, and it does so at the moment the first virtual user trips the condition. The run then covers a shorter window than you declared and is not comparable with earlier runs. Deciding whether a completed run passed is a separate mechanism with its own place in the simulation.
- What does dropping users with exitHereIfFailed do to the load you actually applied?It removes their remaining requests, so every step after the guard sees fewer users than the injection profile describes, and nothing in the report announces the shortfall. That is usually the right trade — meaningless requests are worse than missing ones — but the drop rate is something you have to look at deliberately.
- When is crashLoadGenerator preferable to stopLoadGenerator?When a run that quietly ended early would be worse than one that is unmistakably marked as having gone wrong — for example a guard that fires because the target environment is not the one you meant to load. The graceful flavour ends the run cleanly; the crash flavour ends it as a failure.
A retry is a substitution during play, an exit is sending one player off, and stopping the load generator is abandoning the match. All three are legal; only one of them means nobody gets a result.
saying these in an interview costs you the question
- Using stopLoadGenerator as the run's pass or fail gate
- Putting a blanket tryMax around every step in a scenario
- Expecting exitHereIfFailed to break only the enclosing loop
- Retrying a non-idempotent call such as an order submission
- Treating a dropped user as free, with no effect on applied load