Across a Gatling suite, would you declare every steady arrival phase as constantUsersPerSec(rate).during(d) or append .randomized(), and what does Gatling actually let you control about that choice?
answer
- Two settings, nothing in between
- Plain is deterministic, randomized is not
- No seed, no jitter percentage
- Decide per phase, write it down
basics
~20 sThere is no universally right answer, but Gatling fixes the terms: the plain step reruns with an identical arrival schedule, while the randomized one takes its seed from the clock and can never be replayed.
solid answer
~40 sGatling offers exactly two settings and no dial between them. The plain step produces an identical arrival schedule on every run, so any difference between two runs cannot have come from the injector. `.randomized()` produces irregular, clustered arrivals at the same average rate — but its seed is taken from the system clock and the DSL exposes no way to pin it, so that schedule can never be reproduced. My default is the plain step for phases whose results get compared run to run, and `.randomized()` only where clustering is the thing being exercised. Whichever you pick, make it a written convention: the two declarations differ by a single appended method call, and the source file is the only place that records which one a run used.
go deeper
Be ready to recall that both declarations hold the same average rate over the same window, and that what changes is the spacing between user starts.
Be ready to explain that the plain step reruns identically while the randomized one takes its seed from the clock and cannot be replayed.
Be ready to say which phases of a real profile you would leave deterministic, and how you would handle a finding that only a randomized run produced.
Be ready to set the convention for a whole suite, defend it against the realism argument, and name what Gatling refuses to give you to make it negotiable.
## What Gatling actually puts on the table This is a narrow choice with sharp edges, and it helps to state exactly what the tool gives you before arguing about which to pick. | | `constantUsersPerSec(rate).during(d)` | the same with `.randomized()` | |---|---|---| | arrival spacing | even, computed once before the run | drawn from a Poisson process | | same schedule on a rerun? | **yes**, identical offsets | **no**, never | | seed control | not applicable | **none** — taken from the clock, no DSL parameter | | head-count | rate × window in whole seconds, **rounded** | the same product **truncated**, on 3.15.1 and later — one fewer when its fraction reaches a half | | granularity between the two | none. There is no jitter percentage and no partial randomization | The head-count row is close enough to identical to plan around, but it is not literally identical: the plain step rounds `rate × seconds` while the randomized replacement truncates it. They agree on every whole product and wherever the leftover fraction is under a half; where it reaches a half they differ by one user, so `constantUsersPerSec(0.4978).during(100)` starts 50 plain and 49 randomized. Never more than one, and never at a whole rate. The last row is the important one. Gatling does not offer a middle setting. You are choosing between two fixed behaviours, not tuning one. ## The case for the plain step as the default The plain step's arrival schedule is deterministic. Run the same simulation twice against the same build and the users start at exactly the same offsets. That property is worth something concrete: when two runs disagree, you can rule the injector out without argument. Nobody gets to say *"maybe the arrivals clumped differently this time."* It also makes a profile reviewable. A reviewer reading `constantUsersPerSec(100).during(Duration.ofMinutes(10))` can compute exactly what it will do — 60,000 users, one every 10 ms — and check it against what the author claimed. There is no distribution to reason about. ## The case for `.randomized()` The honest argument for the modifier is that the even schedule is an artefact of the tool. Nothing outside a load generator starts work on a perfect 10 ms cadence. If what you are exercising is specifically how the system behaves when arrivals bunch up — a queue, a lock, a batching layer, a connection pool that is fine on average and not fine on a burst — then the even schedule can hide the very effect you came to find, and `.randomized()` is the one lever Gatling offers. What that costs you is repeatability. Because the seed comes from the clock, a run that found something interesting cannot be replayed. You cannot bisect it, you cannot hand the exact schedule to a colleague, and you cannot re-run it after a fix and know the arrivals matched. ## A policy that survives contact A workable rule, stated per phase rather than per suite: 1. **Phases whose numbers are compared across runs** use the plain step. Determinism is the point; give up realism there deliberately. 2. **Phases that exist to provoke burst behaviour** use `.randomized()`, and their findings are treated as leads to reproduce another way, never as a measurement to trend. 3. **Ramps** follow the phase they lead into, so the profile does not change character at the join. 4. **The choice is written down in the repository**, next to the profile. This matters more than which way you decide, because the two declarations differ by one appended method call and the simulation source is the only place the choice is stated. 5. **Nobody adds `.randomized()` to make a profile look more realistic** without saying which behaviour they expect it to expose. "More realistic" is not a result. ## What this choice is not Two boundaries are worth naming so the discussion stays useful. It is **not** a way to model your real traffic. Choosing an arrival pattern that reflects how work actually reaches your system is a workload-modelling question, and Gatling's contribution to it is only these two declarations. Do not let "we turned on randomized" stand in for having thought about the shape of the load. And it is **not** a tuning knob. Because there is no seed, no percentage and no partial form, there is no incremental version of this decision. You cannot dial the irregularity up for a soak and down for a regression run — you can only pick one per step. ## How to defend the decision The strongest answer in an interview is not "always use randomized" or "always use the plain form." It is: *the plain step is the default because determinism is cheap and repeatability is expensive to lose; the modifier is applied deliberately, to named phases, for a stated reason, and its findings are re-established on a deterministic profile before anyone acts on them.* That answer shows you know both what the tool guarantees and what it refuses to give you.
- A run using .randomized() surfaced a latency spike. What do you do with it?Treat it as a lead, not a measurement. The schedule that produced it cannot be replayed, so reproduce the effect on a deterministic profile — a higher flat rate, or a shorter window — before anyone sizes a fix against it or puts a number in a report.
- Someone proposes making .randomized() the suite-wide default. What is your objection?That it removes repeatability from every phase at once, including the ones whose results get compared across runs. If arrivals differ on every run, a change between two runs can always be blamed on the injector, and that argument is unfalsifiable.
- Can you compromise by randomizing only part of a phase?Not within one step. The modifier applies to the whole step and has no partial form. You can split a phase into two consecutive steps and randomize only one of them, but that is a different profile shape, not a softer setting.
saying these in an interview costs you the question
- Claiming randomized arrivals make a profile automatically more realistic
- Expecting a jitter percentage or seed to tune the choice
- Trending numbers from a run whose schedule cannot be replayed
- Leaving the choice implicit in each simulation file
- Believing the modifier also raises the average arrival rate