skip to content

A Gatling login simulation feeds from a 200,000-row credentials CSV. Does Gatling hold every row in heap, and what decides?

level: seniorimportance: should knowfreq 36%

answer

  1. small file in heap, big file streamed
  2. Gatling picks the loading mode itself
  3. size compared against a configured threshold
  4. feederAdaptiveLoadModeThreshold, 100 MB default
  5. json, sitemap and jdbc always load fully

basics

~20 s

It depends on file size. Gatling compares the uncompressed file length against gatling.core.feederAdaptiveLoadModeThreshold, 100 MB by default: below it the whole file is parsed into memory, above it records are read from disk as users consume them.

solid answer

~50 s

For the character-separated feeders — `csv`, `tsv`, `ssv`, `separatedValues` — Gatling chooses per file, at run start. It measures the file's length and compares it with `gatling.core.feederAdaptiveLoadModeThreshold` from `gatling.conf`, which defaults to **100** megabytes. Under the threshold it parses the whole file into an in-memory vector; over it, it keeps a channel open and reads records in batches. A 200,000-row credentials file of a couple of megabytes is therefore fully in heap, which is fine and needs no tuning. The choice has been automatic by default since 3.1; what changed in 3.15 is that the `eager` and `batch` modes that used to let you force it were removed, leaving the threshold as the only control. Two caveats: the threshold is measured *after* `unzip`, so a small gzip of a huge CSV still streams; and it applies only to the line-based feeders — `jsonFile`, `sitemap` and `jdbcFeeder` always load their whole source into memory.

code

hocon · 7 lines
hocon
gatling {
  core {
    # File size in MB. At or below this a file-based feeder is parsed
    # into memory; above it records are read from the file as needed.
    feederAdaptiveLoadModeThreshold = 250
  }
}

go deeper

for a junior

Know that you do not configure how a feeder file is loaded. Gatling decides from the file size, and an ordinary data file of a few megabytes is simply held in memory.

for a middle

Explain the threshold by name, its 100 MB default and its unit, and that the decision is made at run start from the file's length rather than from its row count.

for a senior

Reason about a generator's heap: know which feeders bypass the adaptive rule and always load whole, and that unzip makes the uncompressed size the one that counts.

for a principal

Set the estate-wide expectation for how big a feeder file may get and in which format, so nobody discovers the difference between a large CSV and a large JSON array during a soak run.

## Gatling chooses the loading mode for you, per file, at run start For the line-based feeders — `csv`, `tsv`, `ssv`, `separatedValues` — the builder does not read anything when you declare it. It resolves the path and stops. The decision that matters happens when the run starts and Gatling turns the builder into an actual feeder. At that point it measures the file's length on disk and compares it against one configuration key: ```hocon gatling { core { feederAdaptiveLoadModeThreshold = 100 # megabytes } } ``` * **Below the threshold**, Gatling opens a channel, parses the entire file, and keeps the records in an in-memory vector. Every virtual user then draws from heap. * **Above it**, Gatling keeps a `FileChannel` open and reads records from disk as users consume them, holding only a working buffer. The comparison is strictly greater-than — `file.length > threshold` — so a file sitting exactly on the threshold is still loaded whole. The default is **100 MB**, stored in `gatling-defaults.conf` and multiplied by 1,048,576 internally, so the key really is in megabytes. A 200,000-row credentials file of two columns is a handful of megabytes: it lands well under the threshold, is loaded whole, and needs no tuning at all. The question only becomes live in the hundreds of megabytes. ## Which feeders never get the choice The adaptive rule is specific to the line-based sources. Several feeders bypass it entirely, and they bypass it in the expensive direction: | feeder | when the source is read | what is held | |---|---|---| | `csv` / `tsv` / `ssv` / `separatedValues` | at run start | in memory, or streamed above the threshold | | `jsonFile` | at run start | **always fully in memory** | | `jsonUrl` | downloaded to a temp file **when declared** | always fully in memory | | `sitemap` | parsed **when declared** | always fully in memory | | `jdbcFeeder` | query runs **when declared** | whole result set, always in memory | | `redisFeeder` | one command per record, during the run | nothing | So a 400 MB CSV streams and a 400 MB JSON file does not. If your data is large enough for the threshold to matter, the format you chose matters more than the threshold does. ## unzip changes what gets measured `unzip()` is applied **before** the size check. Gatling detects the archive by reading the first two bytes — the `PK` marker for ZIP, `0x1f 0x8b` for gzip — not by file extension, decompresses the source into a temporary file, and only then measures. The consequence is worth internalising: a 20 MB gzip of a 400 MB CSV is **streamed**, because the threshold sees 400 MB. The compressed size you committed is irrelevant. A ZIP archive must hold exactly one entry; more than one raises *ZIP Archive contains more than one file*, an empty one raises *ZIP Archive is empty*, and anything whose magic bytes match neither format raises *Archive format not supported*. The Java API's own javadoc on `unzip()` says "zip or tar" — tar is **not** supported, and a tar file hits that last error. ## What was removed, and when The adaptive choice itself is old: it became the default in **3.1**. What you could still do before 3.15 was override it by hand with `eager` and `batch` on a file-based feeder. Both overrides were **removed in 3.15**: the upgrade note says the automatic behaviour "has been given ample satisfaction", so it became the only behaviour. Two consequences: 1. Code carrying `.batch()` or `.eager()` no longer compiles, and there is nothing to replace them with other than the threshold. 2. Gatling's own cheat-sheet page still documents `batch` with a `(bufferSize)` signature. That page is marked as a draft and is stale — do not take it as current. ## Two things the threshold does not do * It does not cap anything. Nothing is rejected for being too large; the key picks a strategy, not a limit. * It does not watch the file. The measurement happens once, when the run builds the feeder, so a file that grows while the run is in flight is neither re-measured nor re-read from the start. ## What to actually do about a 200,000-row credentials file 1. Leave the loading mode alone. At that row count the file is loaded whole, and a few megabytes of heap is not a load generator's problem. 2. Watch the **format** rather than the size. If the same data were a JSON array it would be held in memory however large it grew. 3. If you ship it compressed, size your expectations from the *uncompressed* file. 4. If you do raise the threshold, do it in `gatling.conf` on the classpath, or as a `-Dgatling.core.feederAdaptiveLoadModeThreshold=` system property, since system properties win over `gatling.conf`, which in turn wins over the defaults. The one thing worth knowing about streamed mode is that it is not a free upgrade of the in-memory one: it reads forward through the file, so any behaviour that needs the whole record set at once is served from a rolling buffer rather than from the file as a whole.

  • Does calling unzip() change which loading mode a Gatling CSV feeder gets?
    It can. Gatling decompresses the source into a temporary file first and measures *that*, so the threshold always sees the uncompressed size. A 20 MB gzip of a 400 MB CSV is read from disk in batches, not loaded whole, even though the file you committed is small.
  • Which archive formats does Gatling's unzip() accept?
    ZIP and GZIP, detected by reading the first two bytes — the `PK` marker or gzip's 0x1f 0x8b — rather than by file extension. A ZIP holding more than one entry is rejected with *ZIP Archive contains more than one file*, and anything else with *Archive format not supported*. The Java API javadoc says "zip or tar"; tar is not supported.
  • Where does gatling.core.feederAdaptiveLoadModeThreshold go, and in what unit?
    In `gatling.conf` on the classpath, or as a `-Dgatling.core.feederAdaptiveLoadModeThreshold=` system property, which wins over the file. The value is in megabytes and Gatling multiplies it by 1,048,576 internally. Its default of 100 lives in `gatling-defaults.conf`.

It is the difference between memorising a shopping list and keeping your finger in a phone book. Short enough and you carry it in your head; long enough and you read on as you go.

saying these in an interview costs you the question

  • Believing you can still force eager or batch loading; both were removed in 3.15.
  • Assuming jsonFile streams a large file the way csv does.
  • Thinking the threshold measures the compressed size of a zipped feeder.
  • Expecting a 200,000-row CSV to need tuning; the default is 100 MB.