In a Gatling closed injection profile, what actually brings the number of concurrent users down when a step ramps from 300 to 50?
answer
- Down is not symmetrical with up
- Only completion removes a virtual user
- The load generator tops up, never tears down
- A short wind-down window is decoration
- The curve trails only a ramp it cannot keep up with
basics
~20 sOnly virtual users finishing their scenario. Gatling never interrupts a running user to meet a lower target: it withholds the replacement for a finished user while the population is still above the target for the current second, so concurrency falls at whichever is slower — the rate scenarios end, or the declared slope.
solid answer
~60 sGatling's own reference is explicit: ramping down the number of concurrent users will not force existing users to interrupt, and the only way a virtual user terminates is by completing its scenario. Mechanically, the load generator re-evaluates the step's target about once a second, starts users when the population is **below** target, and does nothing at all when it is above. Separately, each time a user finishes, a replacement starts only if concurrency is still under the target. During a ramp-down that target is a *falling line*, not the ramp's end value, so the check fails while the population is still above the line and succeeds as soon as completions pull it below — which is how the measured curve comes to track a ramp the scenario can keep up with. Only when completions are too slow to reach the falling target does the check keep failing and the population simply shrink one user at a time: if a scenario iteration takes three minutes, a two-minute ramp from 300 to 50 will still be near 300 when the window ends.
go deeper
Be ready to quote the rule: a ramp down does not interrupt anyone, and a virtual user ends only by completing its scenario.
Be ready to explain the correction loop — the target is re-evaluated each second, users are added when the population is below it and nothing happens when it is above — and to add that a finished user is replaced at once whenever concurrency has dipped under the current target, which is what keeps a ramp-down on its declared line.
Be ready to read a concurrency chart that lags its profile and say confidently that the scenario's iteration length, not the injector, is the cause.
Be ready to set the rule a team writes profiles by: wind-down windows sized in scenario iterations, and a run-length bound that survives a degraded system under test.
## The rule, stated the way Gatling's own documentation states it Gatling's injection reference carries a warning next to the closed model: *"Ramping down the number of concurrent users won't force the existing users to interrupt. The only way for virtual users to terminate is to complete their scenario."* That sentence is the whole subject. A closed injection step declares a **target**, and Gatling only ever moves toward that target in one direction — **up**. ## The correction loop, measured A closed population is driven by a tick. Once per second the load generator: 1. works out which injection step the run is currently inside, from the accumulated step durations; 2. evaluates that step's target for the current moment — `300` for a hold, an interpolated integer for a ramp; 3. subtracts the number of users currently in flight; 4. if the difference is **positive**, starts that many new users, spread across the coming second; 5. if the difference is **negative**, does **nothing at all**. There is no branch that stops a user. Separately, whenever a virtual user finishes its scenario, Gatling checks the current target again and starts one replacement immediately if concurrency is below it — which is what makes the hold a hold. During a ramp-down that check fails while the population is still above the ramp's current interpolated target, and starts succeeding again the moment completions pull it below — so a ramp the scenario can keep up with is held on its declared line, and only a ramp steeper than the completion rate leaves the population trailing above it, shrinking one user at a time. It is worth being precise about which of the two mechanisms does the work on the way down. During a monotone ramp-down the tick above never fires step 4: the immediate replacement on each completion tops the population back up the instant it dips under the falling line, so the tick always finds the population at or above the target and its difference is never positive. The per-completion replacement is the whole story for a ramp-down; the tick's job is reaching a level and following a ramp **upward**. ## What that means for 300 down to 50 Take `rampConcurrentUsers(300).to(50).during(Duration.ofMinutes(2))` after a ten-minute hold. The *declared* target falls by roughly two users per second. The *measured* concurrency falls at whichever of those two is slower — that declared slope, or the rate scenarios end: | scenario iteration length | what the concurrent-user chart shows over the two minutes | |---|---| | a few seconds | the measured curve tracks the declared ramp closely | | about one minute | the curve lags, then catches up near the end of the window | | three minutes | almost nothing happens; the run is still near 300 when the ramp ends | So the profile you wrote and the curve the run draws can be two different lines, and the size of the gap is a property of your **scenario**, not of your injection profile: the measured decay is the slower of the declared slope and the rate scenarios complete, and the gap opens only when completion is the slower of the two. This is the single most common surprise in Gatling's closed model, and it is why a wind-down window shorter than one scenario iteration is decoration. ## Why Gatling refuses to interrupt A virtual user is in the middle of something: an HTTP request in flight, a pause, a loop iteration, a group that has not closed. Cutting it off would produce a request with no response, a half-written group timing and a session that never reached its exit hook — measurements that would then pollute the statistics the run is there to produce. Letting the scenario finish keeps every recorded request attributable to a complete journey. The trade is predictability of *timing* for integrity of *data*, and Gatling chose the data. ## The same rule at the end of the run The end of the last step is a ramp-down to zero in disguise, and it behaves identically. When the profile's clock runs out, Gatling stops topping up and marks the population as fully scheduled, but the run does not end until every user it already started has finished. A ten-minute profile whose scenario takes three minutes is a thirteen-minute run at worst. If the wall clock matters — a CI job with a timeout, for example — bound it with `maxDuration` on `setUp`, which ends the run gracefully rather than crashing it. ## Making the curve match the profile There is no interrupt switch in the closed injection DSL, so every remedy works on the other side of the equation: - **Make the wind-down window longer than one scenario iteration.** If an iteration takes a minute, a five-minute ramp from 300 to 50 will track; a thirty-second one will not. - **Shorten the scenario** — fewer actions per iteration, or a loop body that ends sooner — so users turn over more often and the population can respond faster. - **Accept the trail and say so in the report.** For many runs the wind-down is not the part being measured, and a curve that lags for a minute costs nothing. - **Do not expect a steeper ramp to fix it.** The slope is a real lever, but only up to the completion-rate ceiling: once the target falls faster than scenarios end, completion is the binding constraint and steepening buys nothing. In the three-minute-iteration case above it is already binding, which is exactly why the curve sits at 300. ## Common mistakes - Reading a lagging concurrency curve as a broken load generator or a saturated injector. - Declaring a wind-down shorter than a single scenario iteration and expecting it to be honoured. - Assuming the profile's total declared duration is the run's duration. - Believing a lower target eventually forces users out if you wait inside the same step — it never does; only completion removes a user.
- How do you make a Gatling closed-model run actually reach 50 concurrent users inside the window you declared?Work on the scenario, not the profile. Make the wind-down window at least as long as one full scenario iteration, or shorten the iteration so users turn over faster. The closed injection DSL has no interrupt switch, and once the declared slope is already steeper than the rate scenarios finish, steepening it further buys nothing — completion is the binding constraint from there on.
- Does the same rule apply at the very end of a closed profile?Yes. When the last step's window expires Gatling stops starting replacements and marks the population fully scheduled, but the run ends only once every started user has finished. Bound the wall clock with `maxDuration` on `setUp`, which ends the run gracefully rather than crashing it.
- Why does Gatling refuse to interrupt a user rather than offering an option to do so?An interrupted user leaves a request with no response, an unclosed group timing and a session that never reached its exit hook — all of which would pollute the statistics the run exists to produce. Letting the scenario finish keeps every recorded request attributable to a complete journey.
A held concurrent count is like a car park with a fixed number of passes in circulation. Lowering the sign from 300 passes to 50 tows nobody out; it only withholds a new pass as each car leaves, and only while more than 50 are still inside. Lower the sign a notch at a time instead and the attendant keeps handing passes back — the sign, not the gate, then sets the pace, as long as cars leave faster than the sign falls.
saying these in an interview costs you the question
- Believing a lower target forces surplus virtual users to stop.
- Declaring a wind-down shorter than one scenario iteration.
- Reading a lagging concurrency curve as a broken load generator.
- Treating the declared profile duration as the run's duration.