In Gatling, why does a feeder declared as csv("users.csv").shard() still hand every virtual user rows from the whole file on a developer machine?
answer
- an Enterprise-only feeder option
- inert on a single machine
- slices a file across load generators
- opt-in, unlike population sharding
basics
~20 sBecause shard only slices a file when Gatling Enterprise distributes the run across load generators. Everywhere else it is a documented no-op: the option is set on the feeder and nothing in the run path reads it.
solid answer
~50 s`shard` marks a feeder as one whose records should be divided among the load generators of a distributed run, so a 30,000-row file across three generators becomes a 10,000-row slice each. The method is declared on the base feeder builder rather than on the file-based subtype, so it also compiles on in-memory and JDBC feeders, even though Gatling documents and illustrates it for data files. Only Gatling Enterprise performs that division. In the Community Edition the call is documented as a no-op: it flips a boolean on the feeder's options that no local feeder source ever reads, so the feeder behaves exactly as it would without it and every virtual user draws from the whole file. Note the polarity against the population-level `noShard`: a feeder shards only if you ask, since the option defaults to off, whereas a population is sharded by default and `noShard` opts it out. Java, Kotlin and JavaScript write `.shard()`; Scala writes `.shard`.
code
java · 19 linesimport static io.gatling.javaapi.core.CoreDsl.*;
import static io.gatling.javaapi.http.HttpDsl.*;
import io.gatling.javaapi.core.FeederBuilder;
import io.gatling.javaapi.core.ScenarioBuilder;
import io.gatling.javaapi.core.Simulation;
public class ShardedFeederSimulation extends Simulation {
private final FeederBuilder<String> users = csv("users.csv").shard();
private final ScenarioBuilder scn =
scenario("login").feed(users).exec(http("login").post("/login"));
{
setUp(scn.injectOpen(atOnceUsers(100)))
.protocols(http.baseUrl("https://example.com"));
}
}go deeper
Recall that shard is an Enterprise-only feeder option that does nothing on one machine. Be ready to say it slices a file across load generators rather than changing how records are consumed.
Be ready to explain the two defaults side by side: feeder sharding is opt-in, population sharding under Gatling Enterprise is on by default and opted out with noShard. Two words, opposite polarities.
Be ready to say how you keep several hand-run injectors off each other's data when shard cannot help, and why the answer has to be visible in the run rather than implied by a call that does nothing.
Own the risk of shipping an option that is inert in the edition you actually run: a call that does nothing locally and changes behaviour on Enterprise is a correctness difference between environments, not a convenience.
## What the option is for `shard` is a **feeder option** that marks a feeder's records as ones to be divided among the load generators of a **distributed run**. Gatling's documentation gives the arithmetic, under *Distributed files*: a file with 30,000 records deployed on 3 load generators means each generator uses a 10,000-record slice. That is what stops three machines from logging in as the same 10,000 users. Be precise about the **scope**, because the name and the documentation both point at files while the API does not. The method is declared on the **base** feeder builder — `FeederBuilder<T>.shard()` in the Java API that Java and Kotlin share, `FeederBuilderBase[T].shard` in Scala — and not on the file-based sub-interface, so `listFeeder(data).shard()` on an in-memory feeder and `jdbcFeeder(...).shard` on a JDBC one compile perfectly well. Contrast `unzip`, which genuinely is restricted to the file-based subtype. What is file-specific here is the documentation and the worked example, not the type; and off Gatling Enterprise every one of these is equally inert. Attaching it looks the same as any other feeder option: - Java, Kotlin, JavaScript and TypeScript: `csv("users.csv").shard()` - Scala: `csv("users.csv").shard` ## Why it does nothing on your machine The division is performed by **Gatling Enterprise**, not by the feeder. Gatling's documentation is blunt about it — the option is *"only effective when running with Gatling Enterprise Edition, otherwise it's just a noop"* — and the Java API repeats the same sentence in the method's own contract. Mechanically, the call sets a boolean on the feeder's options object. In the Community Edition the feeder sources that build the actual record iterator read the other options — the consumption strategy, the record transform, whether the file needs unzipping — and never read that boolean. It is written and never consulted. So the feeder behaves **exactly as it would have without the call**: every virtual user on that machine draws from the whole file, in whatever order the consumption strategy dictates. ## Three things it is not | people expect | what actually happens | |---|---| | deduplication, so no two users get one record | it never touches per-user allocation on one machine | | a change of consumption strategy | strategy is chosen separately; `shard` does not alter it | | a local slice of the file | there is no node count locally to slice by | The first of these is the costly one. A team that adds `.shard()` because two virtual users kept colliding on the same account will see no change, conclude the feeder is broken, and start debugging the wrong thing. ## The polarity trap: `shard` versus `noShard` Gatling has two sharding words and they point in opposite directions: | | feeder `shard` | population `noShard` | |---|---|---| | what it names | a feeder's records | an injected population | | default | **not** sharded | **is** sharded, on Enterprise | | the call's job | opt **in** to slicing | opt **out** of slicing | | off Enterprise | documented no-op | documented no-op | Read the pair once and it stays straight: you ask for your **data** to be divided, and you ask for your **load** not to be. ## What to do instead when you are hand-running several injectors If you are running Gatling's Community Edition on several machines yourself, `shard` will not separate their data and there is no other built-in that will. The separation has to come from outside the feeder: 1. **Give each machine its own file.** Split the data ahead of time and point each injector at its own path. Blunt, obvious in the artefacts, and impossible to get subtly wrong. 2. **Compute a slice in the simulation.** Read a node index and a node count from a system property or an environment variable, and take that slice of the records yourself before feeding them. `deploymentInfo` cannot supply those numbers off Gatling Enterprise — its generator index is hardwired to 0 and its generator count to 1 — so the numbers have to come from you. 3. **Make the data generated rather than shared.** Where the scenario allows it, deriving values per virtual user removes the collision problem instead of dividing it. Whichever you choose, the important part is that it is **visible**. A `.shard()` call that quietly does nothing looks like the problem is handled; a per-machine file path or an explicit node index does not. ## The version note worth carrying `shard` is current and unchanged on the 3.15.x line. It is worth saying so explicitly, because the neighbouring feeder **loading** modes were not so lucky: the `eager` and `batch` loading options were dropped in 3.15 in favour of an automatic threshold. `shard` is a different option, it is not deprecated, and it still means what it has always meant — just not on a single machine.
- Does Gatling's feeder shard option change which record a given virtual user receives on a single machine?No. On one machine it is inert: the feeder yields records exactly as it would without the call, in whatever order its consumption strategy dictates. Only a distributed Gatling Enterprise run turns the flag into a real per-generator slice of the file.
- Gatling's population noShard and the feeder shard option both concern distribution. How do their defaults differ?They are opposites. A population is sharded by default under Gatling Enterprise and `noShard` opts it out; a feeder is not sharded by default and `.shard()` opts it in. Both are documented as no-ops outside Gatling Enterprise.
saying these in an interview costs you the question
- Expecting shard to deduplicate rows within a single Gatling run
- Thinking shard replaces the feeder's record consumption strategy
- Believing shard makes a local run read only part of the file
- Confusing the feeder shard option with a population's noShard