skip to content

Gatling's reference says constantUsersPerSec starts users at regular intervals - at what granularity does Gatling actually build that schedule, and when do the gaps stop being equal?

level: seniorimportance: should knowfreq 24%

answer

  1. The rate is resolved once, upfront
  2. Whole seconds, then whole milliseconds
  3. No fractional users in a second
  4. A sub-second window starts nobody
  5. Flat rate is a ramp underneath

basics

~20 s

Gatling resolves the rate into one whole head-count, shards it across whole seconds, then spreads each second's share over that second's 1000 millisecond slots. Gaps are exactly equal only when every second carries the same share and it divides 1000.

solid answer

~50 s

The rate is not fed to a live scheduler. `constantUsersPerSec(rate).during(d)` first resolves a single whole head-count — the rate times the window in **whole seconds**, rounded — and then hands it to the same code path a count-based ramp uses. That schedule is built in two stages: the head-count is sharded across whole seconds, and each second's whole-number share is then spread over that second's 1000 millisecond slots. So the spacing is exactly uniform only when both stages come out even: every second must carry the **same** share, and that share must divide 1000. `constantUsersPerSec(20)` qualifies at one arrival every 50 ms; `constantUsersPerSec(3)` does not, and runs 333, 333, 334 ms. `constantUsersPerSec(2.5).during(4)` starts 10 users at 3, 2, 3, 2 a second, with gaps swinging between 333 ms and 500 ms. A window under a second truncates to zero whole seconds and starts nobody, while still consuming its wall-clock time.

code

scala · 20 lines
scala
import io.gatling.core.Predef._
import io.gatling.http.Predef._

import scala.concurrent.duration._

class GranularitySimulation extends Simulation {

  private val httpProtocol = http.baseUrl("https://example.com")

  private val scn = scenario("granularity").exec(http("home").get("/"))

  setUp(
    scn.inject(
      // 10 users: per-second counts are 3, 2, 3, 2 - not a flat 2.5
      constantUsersPerSec(2.5).during(4.seconds),
      // 0 users: 500 ms truncates to 0 whole seconds, but the step still costs 500 ms
      constantUsersPerSec(20).during(500.milliseconds)
    )
  ).protocols(httpProtocol)
}

go deeper

for a junior

Be ready to recall that the schedule is worked out before the run starts and that a second always carries a whole number of users.

for a middle

Be ready to explain the two sharding stages — whole seconds, then milliseconds within each second — and why a flat rate step is a count-based ramp underneath.

for a senior

Be ready to diagnose a profile whose phase silently started nobody because its window was under a second, and to recognise low-rate spacing wobble as scheduler granularity.

for a principal

Be ready to say where this granularity stops being adequate for what you are trying to declare, and what you would do about it rather than fighting the DSL.

## The rate never survives as a rate It is tempting to picture `constantUsersPerSec(20).during(15)` as a token bucket ticking away during the run. It is not. Gatling resolves the whole arrival schedule **up front**, as a sequence of offsets from the start of the simulation, and the rate is consumed in the first line of that resolution. The step computes a single head-count: the declared rate multiplied by the window expressed in **whole seconds**, rounded to a whole number of users. It then delegates to exactly the same machinery a count-based ramp uses. Gatling's Scala source comments the flat rate step as *"Inject users at constant rate : another expression of a RampInjection"*, and the Java API's javadoc on the count-based ramp builder returns the compliment, describing it as *"Strictly equivalent to ConstantRate"*. `constantUsersPerSec(20).during(15)` and `rampUsers(300).during(15)` produce the identical schedule, because after the first line they *are* the same object. ## Two stages of sharding The schedule is built in two nested stages, both over integers: 1. **Whole seconds.** The window is truncated with a seconds conversion, and the head-count is sharded across that many whole seconds. Each second gets a whole number of users; there is no such thing as half a user in a second. 2. **Milliseconds.** Each second's share is then sharded across that second's 1000 millisecond slots, and users are emitted at the slot boundaries. Both stages use the same integer sharding helper, which walks a running ceiling so the shares sum exactly to the total and the first bucket is never empty. ## When the gaps are equal, and when they are not Exactly equal gaps — one arithmetic progression across the whole window — need **both** stages to come out even: every second must carry the **same** whole-number share `k`, and `k` must **divide 1000**. Among whole rates only the divisors of 1000 qualify — 1, 2, 4, 5, 8, 10, 20, 25, 40, 50, 100, 125, 200, 250, 500, 1000: * `constantUsersPerSec(20).during(15)` → 300 users, 20 in every second, one arrival every **50 ms**, uniform throughout. * `constantUsersPerSec(100).during(Duration.ofMinutes(10))` → 60,000 users, 100 a second, one every **10 ms**. "The rate divides cleanly" is *not* the criterion, in either direction. A clean rate of 3 gives every second the identical share of 3, but 3 does not divide 1000: the millisecond shard's running ceiling emits 0, 333, 666, 1000, 1333, … so the gaps run 333, 333, **334**. `constantUsersPerSec(7)` runs 142, 142, **143**. Nobody notices a millisecond — but the rule that predicts it is the one to carry. The wobble you *can* see comes from the first stage, when the per-second shares alternate. Take `constantUsersPerSec(2.5).during(4)`: | second | users in that second | arrival offsets | |---|---|---| | 0 | 3 | 0, 333, 666 ms | | 1 | 2 | 1000, 1500 ms | | 2 | 3 | 2000, 2333, 2666 ms | | 3 | 2 | 3000, 3500 ms | Ten users in four seconds, the declared average of 2.5 a second — but the per-second counts alternate 3, 2, 3, 2, and consecutive gaps run 333, 333, 334, 500, 500, 333, … The mean gap is the promised 400 ms; no individual gap is. `constantUsersPerSec(1.5).during(4)` alternates 2, 1, 2, 1 and runs 500, 500, 1000 ms — even though both of those shares *do* divide 1000. It is the alternation, not the divisor, that breaks it there. A rate below one user a second escapes the same-share rule by carrying zero in most seconds, and can still be perfectly uniform. `constantUsersPerSec(0.5).during(Duration.ofMinutes(1))` resolves to 30 users over 60 seconds and starts one user every *other* second, because a second can never carry half a user — an exact 2000 ms every time, because 2 divides the 60-second window. Shorten the window and that exactness goes: `constantUsersPerSec(0.5).during(5)` rounds to 3 users and runs 1000, 2000 ms. ## Two traps that follow from whole-second truncation Both of these come from the window being converted to whole seconds before anything else happens. * **A sub-second window starts nobody.** `constantUsersPerSec(20).during(Duration.ofMillis(500))` truncates to zero whole seconds, so the head-count is zero and no user is injected — yet the step still consumes its 500 ms of wall-clock time before the next step begins. Nothing warns you. (A literally zero duration is different: that is rejected outright when the step is constructed.) * **A fractional-second window loses its tail.** `constantUsersPerSec(20).during(Duration.ofMillis(2500))` truncates to two whole seconds, so it starts 40 users across the first two seconds and then sits idle for the remaining half-second before moving on. The practical rule: express rate-step windows in whole seconds. There is no resolution below one second in this part of Gatling. ## Fractional rates are fine — the head-count is not fractional None of the above makes fractional rates unsupported. The parameter is a `double`, and Gatling's reference states outright that rates may be expressed as fractional values. What is not fractional is the number of users: the head-count is rounded. Gatling's own test suite asserts that a rate of `0.4978` held for 100 seconds resolves to exactly 50 users. ## Why this matters in practice Two reasons. First, when you compare a report's arrival-rate plot against the profile you declared, small wobbles — a millisecond at rates that miss the divisors of 1000, hundreds where the per-second shares alternate — are the scheduler, not the system under test; there is nothing to investigate. Second, if you need finer control than whole seconds give you, this part of the DSL will not provide it, and reaching for sub-second windows quietly produces empty steps rather than an error.

  • Is constantUsersPerSec(20).during(15) genuinely the same thing as rampUsers(300).during(15)?
    Yes. The flat rate step resolves its head-count first and then delegates to the count-based ramp, so the two produce an identical arrival schedule. Gatling's source comments the rate step as another expression of the ramp, and the Java API javadoc says the two are strictly equivalent.
  • What happens if you declare a rate step with a window of a few hundred milliseconds?
    The window truncates to zero whole seconds, so the head-count is zero and no user starts — but the step still consumes its wall-clock duration before the next one begins. Nothing errors, so the profile silently loses a phase.
  • Can you get sub-second precision on the arrival schedule?
    Not through the window. Arrivals land on millisecond boundaries within a second, but the head-count is derived from whole seconds only, so a window shorter than a second contributes no users. Raise the rate rather than shortening the window.

It works like a bus timetable printed before the day starts, not a driver watching a clock. Each hour gets a whole number of buses spread across it, so asking for two and a half an hour gets you three in one hour and two in the next.

saying these in an interview costs you the question

  • Picturing a live token bucket rather than a precomputed schedule
  • Assuming a rate that divides cleanly always gives equal gaps
  • Assuming a 500-millisecond rate window injects a few users
  • Believing a fractional rate is rejected by the builder
  • Reading low-rate spacing wobble as a defect in the system