A Gatling login simulation needs 200,000 distinct credentials per run. Would you ship them as a CSV, read them with jdbcFeeder, or serve them from Redis?
answer
- who pays for the data, and when
- one of the three works in every SDK
- a query that runs before the run starts
- Redis couples the generator to a live service
- csv keeps the data inside the artifact
basics
~20 sDefault to the CSV. It is the only one of the three available in all five SDKs, it costs the running system nothing, and it ships with the artifact. Use the others only when the data cannot be a file.
solid answer
~40 sThe three differ in what they cost the load generator and when. A CSV under `resources` is resolved at declaration and parsed at run start, reads from the file rather than heap once it is large, and is the only option the JavaScript and TypeScript SDK has. `jdbcFeeder` runs its SQL once, in the Simulation constructor, and materialises the entire result set into memory before the run starts; it needs a JDBC 4 driver on the classpath and a database reachable from every generator. `redisFeeder` is the lazy one — one Redis command per record during the run — but it exposes none of `transform`, `readRecords`, `recordsCount`, `shard` or `unzip`, because its builder is not a feeder builder at all in the Java API. Weigh reproducibility and generator coupling first, convenience second.
code
java · 19 linesimport io.gatling.javaapi.core.*;
import io.gatling.javaapi.redis.*;
import static io.gatling.javaapi.core.CoreDsl.*;
import static io.gatling.javaapi.jdbc.JdbcDsl.*;
import static io.gatling.javaapi.redis.RedisDsl.*;
public class CredentialSources {
// Shipped with the artifact; every chained option still works.
FeederBuilder.FileBased<String> fromFile = csv("data/credentials.csv");
// Query runs here, in the constructor; whole result set held in memory.
FeederBuilder<Object> fromDatabase =
jdbcFeeder("jdbc:postgresql:gatling", "user", "pass", "SELECT username, password FROM load_users");
// One Redis command per record; nothing else chains onto this builder.
RedisClientPool pool = new RedisClientPool("localhost", 6379);
RedisFeederBuilder fromRedis = redisFeeder(pool, "credentials").LPOP();
}go deeper
Know that the file is the normal answer and that the other two exist. Being able to say why a CSV is the default is enough at this level.
Contrast the three on when the reading happens and what the generator holds, and know that jdbcFeeder executes its query at declaration rather than during the run.
Reason about the coupling you take on: a live database or Redis in the request path, a heap you did not size, and a failure mode that appears mid-run instead of at startup.
Own the rule for the estate — whether virtual-user data may come from a live service at all, who owns that service's availability during a run, and how a run stays reproducible months later.
## Ask what each one costs, and when it costs it These three are not three spellings of the same idea. They differ in when the I/O happens, what the load generator holds, what has to be reachable from the generator, and how much of the DSL still works afterwards. | | `csv("data/credentials.csv")` | `jdbcFeeder(url, user, pass, sql)` | `redisFeeder(pool, "credentials")` | |---|---|---|---| | when the source is read | resolved at declaration, parsed at run start | **query runs at declaration** | one command per record, during the run | | what the generator holds | in memory, or read from the file above the size threshold | **the whole result set, always** | nothing | | what must be reachable | nothing at run time | the database, from every generator | Redis, from every generator, for the whole run | | chained options available | `unzip`, `transform`, `readRecords`, `recordsCount`, `shard`, the four consumption strategies | `transform`, `readRecords`, `recordsCount`, `shard`, the strategies — but not `unzip` | **none of them** | | available in the JS/TS SDK | yes | no | no | | extra classpath requirement | none | a JDBC 4 driver jar | the Redis client module | Three of those rows decide most real cases. ## The three facts that do the deciding 1. **`jdbcFeeder` is eager and total.** It opens the connection and executes the SQL while the Simulation constructor is still running, walks the entire `ResultSet` forward, and materialises every row into an in-memory vector before anything else happens. There is no streaming form and the adaptive file threshold does not apply to it. 200,000 rows of credentials is fine; a query that quietly returns ten million is a heap you did not plan for, at startup, on every generator. 2. **`redisFeeder` is not really a feeder builder.** In the Java API its type implements only a supplier of records; it does not implement the feeder-builder interface at all. So `transform`, `readRecords`, `recordsCount`, `shard`, `unzip` and the four consumption strategies are simply absent — not deprecated, absent. The only thing you configure is the Redis command, and the only way to recycle records is to point `RPOPLPUSH` at the same key for source and destination. It is also the only one of the three that puts a network round trip between a virtual user and its next record. 3. **Only the CSV works everywhere.** `jdbcFeeder` and `redisFeeder` are documented as `NOT SUPPORTED` in the JavaScript and TypeScript SDK. If the suite must be portable across SDKs, or might be, the question is already answered. ## Three questions that settle it faster than the table * Does the suite have to run, now or later, in the JavaScript or TypeScript SDK? If yes, it is the file. * Must the records be the ones the target holds at the instant the run starts? If no, it is the file. * Is anything gained by the stock being shared and consumed across processes? If no, it is the file. ## The decision **Default to the file.** For 200,000 credentials, a CSV under `resources` is read once, holds a few megabytes, keeps every chained option, works in all five languages, costs the target system nothing, and reproduces exactly when someone reruns the build months later. The run depends on nothing but the artifact it shipped in. **Reach for `jdbcFeeder`** only when the credentials must be the ones that exist in the target *at this moment* and cannot be exported beforehand — for example when the same pipeline seeds them minutes earlier and the identifiers are not predictable. Accept the price knowingly: an eager query, the whole result set resident, a database credential on every generator, and no JavaScript SDK. **Reach for `redisFeeder`** when the *interesting* property is that the stock is shared and consumed across processes — several generators drawing from one pool without a coordinator, or a producer topping the list up while the run proceeds. That is a genuine capability nothing else here has. Pay for it with a live dependency in the request path and a builder that supports nothing else. ## Where this stops being a Gatling question How those 200,000 credentials are produced, masked, kept distinct, and kept in step with the target's own data is a test-data question, and it is not decided here. Nor is how many virtual users draw on the feeder, which belongs to the injection profile. What Gatling settles is narrower and worth stating precisely in an interview: **which builder reads the data, when it does the reading, and what the load generator is holding while the run is in flight.** A useful tie-breaker when the answer is close: prefer the option whose failure happens before the run rather than during it. A missing or misnamed CSV and an unreachable database both fail in the constructor, loudly, with no report. Redis failing halfway through a two-hour soak fails as a rising error rate that you then have to attribute — and the run is gone either way.
- What does redisFeeder give up compared with a csv feeder in Gatling?Everything chained. Its builder is not a feeder builder at all in the Java API — it implements only a supplier of records — so `transform`, `readRecords`, `recordsCount`, `shard`, `unzip` and the four consumption strategies are simply absent. The only choice you make is the Redis command, and recycling means pointing `RPOPLPUSH` at the same key for source and destination.
- When is jdbcFeeder the right answer in Gatling despite its cost?When the records must be the ones that exist in the target right now and cannot be exported ahead of the run — data seeded by the same pipeline minutes earlier, for instance. Accept that the query runs once in the constructor, that the whole result set sits in heap, and that every generator needs database access and a JDBC 4 driver.
- Does the choice change if the suite might move to Gatling's TypeScript SDK?It settles it. Both `jdbcFeeder` and `redisFeeder` are documented as unsupported there, so a suite that may become JavaScript or TypeScript has to get its virtual-user data from a file under `resources`, from `jsonUrl`, or from an array built in the script.
saying these in an interview costs you the question
- Choosing redisFeeder for recycling without noticing it has no strategy methods.
- Assuming jdbcFeeder streams rows lazily during the run.
- Picking jdbcFeeder for a suite that may also run in TypeScript.
- Treating a database connection from every load generator as free.