skip to content

In a Gatling run, what happens to requests the throttle ceiling will not admit yet, and what ends a run whose throttle profile is shorter than its injection profile?

level: middleimportance: nice to knowfreq 26%

answer

  1. Nothing is dropped, only deferred
  2. An unbounded buffer in the injector
  3. OutOfMemoryError is a generator failure
  4. The profile's length bounds the run

basics

~20 s

Nothing is dropped. Excess requests are buffered in an unbounded in-memory queue inside the load generator and replayed on the next one-second tick, risking an OutOfMemoryError. The throttle profile's total duration also bounds the run, stopping it like maxDuration.

solid answer

~40 s

Gatling's throttler works on a one-second tick: it computes the ceiling for that second, spaces the permitted requests evenly across it, and appends everything over budget to a buffer that is replayed on the next tick. That buffer is unbounded, so a simulation offering far more than the ceiling grows the injector's heap until it throws an `OutOfMemoryError` — a load-generator failure, not a signal from the system under test. Separately, the throttle profile is a duration limit: Gatling sums its `reachRps` and `holdFor` steps, takes the minimum of that and any declared `maxDuration`, and stops the run there. A simulation that never declared `maxDuration` can therefore still end with a *Run stopped on maxDuration(...) reached* message.

code

java · 9 lines
java
setUp(
  browse.injectOpen(constantUsersPerSec(120).during(Duration.ofMinutes(10))),
  checkout.injectOpen(constantUsersPerSec(80).during(Duration.ofMinutes(10)))
).protocols(httpProtocol)
 .maxDuration(Duration.ofMinutes(30))
 .throttle(
   reachRps(200).in(Duration.ofSeconds(30)),
   holdFor(Duration.ofMinutes(5))
 );

go deeper

for a junior

Be ready to say that Gatling defers rather than drops traffic above the ceiling, and that the throttle profile's length also bounds how long the run lasts.

for a middle

Be ready to explain the one-second tick, the replay of the buffer on the next tick, and the minimum taken between maxDuration and the throttle profile's duration.

for a senior

Be ready to diagnose an injector heap failure as a throttle gap, and to say why a throttled run's response times understate the deferral the generator introduced.

for a principal

Be ready to set a rule for when a suite may throttle at all, given that the mechanism trades a load-generator memory risk for a ceiling nothing else enforces.

A Gatling throttle never rejects a request. When the current second's budget is spent, the request is **parked**, and Gatling comes back for it on the next tick. Understanding that queue, and understanding that the throttle profile also fixes when the run ends, are the two things that turn a throttle from a surprise into a tool. ## How the second is spent The throttler wakes on a fixed one-second tick. On each tick it recomputes the current ceiling from the throttle profile, allocates that many permits, and then: * A request that arrives while permits remain is **sent**, and the permit count is incremented. * A request that arrives with no permits left is **appended to a buffer**. * On the next tick the buffer is replayed first against the fresh permits, and anything still over budget is buffered again. Within a tick the permitted requests are not fired as a burst. Gatling divides the second by the current limit and spaces the requests across it, so a ceiling of 200 becomes roughly one request every five milliseconds rather than 200 at the top of the second. ## The buffer is unbounded — and that is the risk That buffer is a plain growable collection with no capacity limit and no drop policy. Gatling's own documentation flags the consequence in a warning: *"all excess traffic gets pushed into an unbounded queue, possibly resulting in an OutOfMemoryError if your normal throughput is way higher than the normal one"*. The failure mode is entirely **inside the load generator's JVM**, not in the system under test: the target never sees the parked requests, so it cannot push back on them, and the injector's heap absorbs the whole excess. The arithmetic is easy to get wrong. A simulation that offers 1,000 requests per second under a ceiling of 200 parks 800 every second. Over a ten-minute hold that is roughly half a million buffered parked request closures, each holding the work still to be done, and every one of them expects to be sent eventually. Nothing sheds, nothing expires, and nothing warns you before the heap goes. Practical consequences: 1. **Size the ceiling near the offered traffic, not far below it.** A throttle is for trimming an overshoot, not for holding back an order of magnitude. 2. **Watch the injector, not just the target.** A throttled run that dies with an `OutOfMemoryError` is reporting a load-generator problem, and the run's results are void. 3. **Do not read a throttled run's response times as the system's.** A request parked for several seconds before it is sent is not timed from when it was parked, so the deferral is invisible in the report. ## The throttle also ends the run The second surprise is that a throttle profile is a **duration limit**. Gatling sums the durations of the profile's `reachRps` ramps and `holdFor` plateaus — `jumpToRps` contributes nothing — and folds that total in with any declared `maxDuration`, taking the **minimum**. When that limit expires the controller stops the run gracefully, with the message *"Run stopped on maxDuration(...) reached"*. Two things follow that regularly confuse people: * A simulation that never wrote `maxDuration` can still stop with a `maxDuration` message, because the throttle supplied the limit. * The limit is computed over the **whole run**, including every population's own throttle. If one population carries a two-minute throttle profile while the others are injecting for an hour, the entire run ends after two minutes. | declared | effective run limit | |---|---| | `maxDuration(30 min)` only | 30 minutes | | `throttle(reachRps(200).in(30), holdFor(5 min))` only | 5 minutes 30 seconds | | both of the above | 5 minutes 30 seconds — the minimum wins | | a 2-minute throttle on one population out of four | 2 minutes, for all four | ## Building a profile that does not surprise you Make the throttle profile's total duration match the run you intended, and let the injection profile be the thing that is generous. In the example below the ceiling ramps for thirty seconds and holds for nine and a half minutes, giving a ten-minute run, while the populations are injected for ten minutes as well — so the throttle bounds the run at precisely the moment the traffic was going to stop anyway, and nothing is cut short.

  • A throttled simulation dies with an OutOfMemoryError. Where do you look first?
    At the gap between offered traffic and the ceiling. Everything over budget is parked in an unbounded buffer inside the injector, so a profile offering far more than the ceiling grows the heap without limit. Bring the ceiling closer to what the profile actually produces, or lower the injection rate.
  • The simulation never calls maxDuration, yet the run stops early with a maxDuration message. Why?
    Because the throttle profile supplied the limit. Gatling sums the profile's reachRps and holdFor durations and treats that total as a run limit, taking the minimum of it and any declared maxDuration. The stop is graceful and the message names maxDuration even though you never wrote one.
  • Does a request parked by the throttler show its waiting time in the report?
    No. The parked request has not been issued yet, so its measurement starts when the throttler finally releases it. Deferral inside the load generator is invisible in the response-time statistics, which is one reason a heavily throttled run's timings should not be read as the system's behaviour.

saying these in an interview costs you the question

  • Assuming excess requests over the ceiling are discarded
  • Reading a throttled run's OutOfMemoryError as a target failure
  • Expecting maxDuration to win over a shorter throttle profile
  • Thinking parked time appears in the reported response times