A Gatling simulation calls csv("credentials.csv").readRecords() only to size its injection profile. What does that cost, and what should it call instead?
answer
- one returns data, one returns a number
- reading the rows is not free
- the reference warns about repeated parsing
- recordsCount counts lines, drops the header
- readRecords is absent from the JavaScript SDK
basics
~20 sreadRecords builds the feeder and drains it into a list, so the file is parsed once for the sizing call and again when the run starts. recordsCount counts lines without constructing records, and is the call for sizing.
solid answer
~60 s`readRecords()` returns the whole source as data — a `List<Map<String, Object>>` in Java, a `Seq[Record[Any]]` in Scala. It works by constructing the configured feeder and draining it, so it genuinely re-reads and re-parses the underlying file; Gatling's reference carries that warning explicitly. Using it just to learn how many rows exist therefore pays for the parse twice: once in the constructor, once again when the run builds the real feeder. `recordsCount()` exists for exactly this case — for a line-based feeder it counts line terminators and subtracts the header, without building a single record. Both are declared on the feeder builder that the file and in-memory sources return — `FeederBuilderBase[T]` in Scala, `FeederBuilder<T>` in the Java API — so `csv`, `tsv`, `ssv`, `separatedValues`, `jsonFile`, `jsonUrl`, `sitemap`, `jdbcFeeder` and `arrayFeeder` all have them. `redisFeeder` is the exception: its builder sits outside that hierarchy and has neither. `readRecords` is also one of the methods the JavaScript and TypeScript SDK does not have. Keep `readRecords` for when you actually want the data.
code
java · 13 linesimport io.gatling.javaapi.core.*;
import static io.gatling.javaapi.core.CoreDsl.*;
public class SizingExample {
FeederBuilder.FileBased<String> credentials = csv("data/credentials.csv");
// Cheap: counts line terminators, builds no records.
int available = credentials.recordsCount();
// Expensive: parses every row into a map, then discards them all.
int alsoAvailable = credentials.readRecords().size();
}go deeper
Know that recordsCount gives you the number of records and readRecords gives you the records themselves, and that you should reach for the first when a count is all you need.
Explain that readRecords re-reads and re-parses the source on every call with no caching, and that recordsCount counts lines and subtracts the header instead.
Spot the double parse in a constructor that sizes an injection from readRecords, and know the recycling-strategy trap that makes draining a feeder non-terminating.
Set the convention for deriving run shape from data: whether the suite reads its own feeder files at startup at all, and what that costs a fleet of generators starting together.
## The two methods answer different questions Both are declared on the feeder builder that the file and in-memory sources return — `FeederBuilderBase[T]` in Scala, `FeederBuilder<T>` in the Java API — so `csv`, `tsv`, `ssv`, `separatedValues`, `jsonFile`, `jsonUrl`, `sitemap`, `jdbcFeeder` and `arrayFeeder` all carry both, and they look interchangeable when all you want is a row count. They are not. * **`readRecords()`** returns the whole source as data — `List<Map<String, Object>>` in Java and Kotlin, `Seq[Record[Any]]` in Scala. It builds the feeder the builder is configured to produce and drains it into a collection. * **`recordsCount()`** returns an `int`. For a line-based feeder it opens the source, counts line terminators using the configured charset, and subtracts one for the header line that CSV, TSV and SSV files carry. Not a single record map is constructed. For a `jsonFile` it walks the array and counts objects. Gatling's own reference introduces `recordsCount` with exactly this motivation: knowing the size of your feeder *without* having to use `readRecords` and copy all the data into memory. ## Why readRecords is the expensive one The reference also carries an explicit warning: each `readRecords` call reads the underlying source again — it parses the CSV file again. There is no cache. So the sequence in the question does this: 1. The constructor calls `readRecords()`. Gatling resolves the file, builds a feeder over it, parses all 200,000 rows into maps, and hands you a list. 2. You call `.size()` on it and throw the list away. 3. The run starts. Gatling builds the real feeder from the same builder and **parses the file a second time**. You paid twice for a number you could have had for the price of counting newlines. On a small file that is invisible; on a large one it is a startup delay and a transient allocation of every record in the file, immediately garbage. Calling `readRecords()` twice compounds it: the second call is not cheaper than the first. ## The trap that turns it from slow into fatal `readRecords` drains **the feeder this builder is configured to produce**. Two of Gatling's four consumption strategies are explicitly documented as never running out of records — they recycle instead. A feeder that never reports exhaustion has nothing to drain to, so draining it does not terminate. The practical rule is simple: call `readRecords()` on the plain builder, before you attach a recycling strategy, or on a separate builder created for the purpose. `recordsCount()` is immune, because it counts the source rather than consuming the feeder. ## Where they exist and where they do not | method | Java / Kotlin / Scala | JavaScript / TypeScript | |---|---|---| | `readRecords()` | yes, on every feeder builder except `redisFeeder` | documented as `NOT SUPPORTED` | | `recordsCount()` | yes, on every feeder builder except `redisFeeder` | yes | That asymmetry is deliberate enough to be worth remembering: a JavaScript simulation can size a feeder and cannot materialise one. Note also that `redisFeeder` has **neither** — its builder is not a feeder builder at all in the Java API, just a supplier of records, so nothing chains onto it. In Scala both are parameterless methods written without parentheses: `csv("x.csv").recordsCount`. ## Neither of them is what feeds the users It is worth being explicit that these two methods are side doors. Calling `readRecords()` does not prime the feeder, and calling `recordsCount()` does not consume anything: the feeder the virtual users draw on is built separately when the run starts, from the same immutable builder. Nothing you read out here is subtracted from what the run gets. ## Sizing an injection from the feeder The usual reason to reach for either of these is wanting the injection profile to match the data: a credentials file whose rows are consumed once, where injecting more users than there are rows ends the run badly. `recordsCount()` is the right call. ```java FeederBuilder.FileBased<String> credentials = csv("data/credentials.csv"); int available = credentials.recordsCount(); setUp(login.injectOpen(atOnceUsers(available)).protocols(httpProtocol)); ``` That reads the file once to count lines and once more at run start to feed, which is the floor. ## When readRecords is the right call It is not a method to avoid — it is a method to use for its actual purpose, which is reusing Gatling's parsers for something Gatling's own feeder does not do: deriving a lookup table, splitting one file into several feeders, or asserting something about the data before the run. It returns plain maps, and an in-memory feeder can be built from a list or array of them. Do it once, keep the value, and remember that the builder is immutable so the value you keep is the only copy you have.
- How does Gatling's recordsCount avoid parsing the records?For a line-based feeder it opens the source and counts line terminators with the configured charset, then subtracts one if the format carries a header, which CSV, TSV and SSV do. For a `jsonFile` it walks the array and counts objects. Either way no record map is ever built.
- Can you use readRecords to build a second Gatling feeder from the same data?Yes, and that is its intended use. It returns plain maps, and an in-memory feeder can be built from a list or array of them, which lets you reuse Gatling's parser for something its own feeder does not cover. Just call it once and keep the value, because every call re-reads the source.
- Why should readRecords be called before a recycling strategy is attached?Because it drains the feeder the builder is configured to produce, and two of Gatling's four consumption strategies are documented as never running out. A feeder that never reports exhaustion has nothing to drain to. `recordsCount` has no such problem: it counts the source rather than consuming the feeder.
saying these in an interview costs you the question
- Calling readRecords().size() when recordsCount() answers the same question.
- Assuming readRecords caches, so that a second call is free.
- Believing recordsCount has to parse every record in order to count.
- Calling readRecords on a feeder configured never to exhaust.