skip to content

Across a Gatling suite, how would you decide which of the four feeder consumption strategies each data file gets, and what can those four not express?

level: principalimportance: should knowfreq 33%

answer

  1. the data decides, not the scenario
  2. uniqueness required means a finite budget
  3. circular sweeps evenly, random draws lumpily
  4. no check-out and return among the four

basics

~20 s

Decide from the data, not the scenario: records carrying an identity get a consuming strategy and a stock sized to the run; everything else gets circular or random. The four cannot express reuse with uniqueness, or lending a record back.

solid answer

~40 s

One question settles it: must the target see each record at most once? Identity-bearing data — logins, one-shot tokens, accounts a scenario mutates — takes `queue` or `shuffle`, and the file then becomes a budget the profile has to stay inside. Everything else takes `circular` or `random`, which cannot run out. Within each pair, `shuffle` over `queue` removes accidental correlation with the file's order, and `circular` over `random` gives an even, predictable sweep rather than a lumpy draw. What the four cannot express is a consuming strategy that wraps, a record lent to one user and returned afterwards, a partition of one file between two scenarios, or any warning as the stock drains — for those you leave the strategy set and supply your own feeder, which is just an iterator of string-keyed maps.

code

java · 10 lines
java
// identity-bearing: must stay consuming, stock sized to the run
FeederBuilder<String> logins = csv("logins.csv").shuffle();

// non-identifying lookups: reuse freely, small file is fine
FeederBuilder<String> terms = csv("search-terms.csv").circular();

// no file at all: a feeder is just an iterator of records
Iterator<Map<String, Object>> orderRefs = Stream.generate(
  () -> Map.<String, Object>of("ref", UUID.randomUUID().toString())
).iterator();

go deeper

for a junior

Be ready to say which strategies reuse records and which do not, and to ask what the data represents before naming one.

for a middle

Explain the two pairs and what distinguishes them within each pair: source order versus permutation, and even sweep versus independent draw.

for a senior

Show the failure mode in both directions — a run that dies for no reason, and a run that finishes while measuring identity collisions your own test data caused.

for a principal

Own the standard: which data classes are consuming, where the strategy is declared, how stock size is provisioned, and when the team drops to a custom feeder instead.

## The decision is about the data, not the scenario Every feeder in a suite answers one question, and the answer picks the strategy: **Must the target see each record at most once, at any moment of the run?** - **Yes** — logins, single-use tokens, coupon codes, accounts the scenario mutates, anything the system under test treats as an identity. Use `queue`, or `shuffle` if you also want the order scrambled. The file is now a budget the profile has to stay inside, and running out stops the run. - **No** — search terms, product identifiers, postcodes, read-only lookups. Use `circular`, or `random` if the order itself matters. The stock can never empty, so an entire class of run failure disappears and the file can be small. Because the answer is a property of the data, the strategy belongs at the **declaration of the feeder**, once, and that declaration should be shared rather than re-typed. Two separate builders over one file do not share a cursor: each walks the file from the top, so a uniqueness guarantee you thought you had quietly stops holding. ## Choosing within each pair - **`queue` vs `shuffle`.** They have identical stock behaviour; only the order differs. `queue` is right when file order is meaningful — records pre-grouped by segment, or a deliberate warm-then-hit ordering. `shuffle` is right when it is not, because it removes an accidental correlation between the order records were generated in and the order the target sees them, which otherwise shows up as a cache or partition artefact you then spend an afternoon explaining. - **`random` vs `circular`.** `circular` sweeps the file evenly and predictably: over a long run every record is used within one of every other. `random` draws independently, so in any finite run the distribution is lumpy — some records go out many times and some not at all. If the data exists to spread load across keys or to defeat a cache, `circular` guarantees the spread and `random` only approximates it. If you want the *order* to be unpredictable — because an ordered sweep is itself an unrealistic access pattern — `random` is the one that gives you that. ## What the four cannot express There are exactly four, they take no parameters, and that bounds what a strategy can do for you: 1. **There is no consuming strategy that wraps.** You cannot say "hand each record out once, then start again". The two strategies that reuse also give up uniqueness, and the two that keep uniqueness also end the run. 2. **There is no check-out / check-in.** A record cannot be lent to one user, made unavailable, and returned to the stock when that user finishes. That is the shape most credential pools actually want, and Gatling does not have it. 3. **There is no partitioning.** Two scenarios in one simulation cannot each be given a disjoint half of one file by strategy alone; they need two feeders over two files. 4. **There is no low-stock signal.** Nothing warns you as the stock drains; the first news is the crash. When you need one of those, you leave the strategy set entirely. A feeder in Gatling is just an iterator of string-keyed maps, so any object with that shape can be handed to `feed` — including one that mints values on the fly and therefore has no stock to run out of. Generating a unique value per user is usually cheaper and more robust than provisioning a file large enough to survive the longest run anyone will ever configure. ## The cost of getting it wrong, in each direction - `queue` on data that did not need to be unique buys you a run that dies for no reason, with no report and no verdict to show for the time spent. - `circular` or `random` on data that did need to be unique buys you something worse: a run that finishes, produces a clean-looking report, and is measuring concurrent sessions colliding on the same identity. The errors and the latency are real; they just belong to the test data. ## A short checklist you can apply per file 1. **What is the record?** An identity the target enforces, or a value it merely reads? 2. **What does the profile draw?** Feed executions over the whole run, not users at peak. 3. **Does order carry meaning?** If the file is grouped or sorted deliberately, keep source order; if the grouping is an accident of how the file was generated, break it. 4. **What happens if this runs out at 3 a.m.?** For a nightly soak, a consuming strategy is a commitment to keep the file ahead of the schedule for as long as the suite lives. Answer those four and the strategy falls out. Skip them and the default decides for you — which is fine for credentials and wrong for almost everything else.

  • Why can switching a credentials feeder to `circular` look like a performance regression?
    Because two virtual users then act as the same principal at once. The target may serialise their sessions, invalidate one another's tokens or reject the second login, and the resulting errors and added latency are real measurements of a collision your test data created — not of the system under load.
  • What do you do when none of the four fits?
    Supply your own feeder. Gatling defines a feeder as an iterator of string-keyed maps, so any object of that shape can be passed to `feed`, including one that generates a fresh unique value per call. That removes the stock entirely and with it the whole class of end-of-data failures.
  • Why prefer `circular` over `random` for cache-busting data?
    Circular walks every record in turn, so the spread across keys is even by construction. Random draws with replacement, so over any finite run some keys are hit repeatedly and others never — which can leave exactly the cache locality the data was chosen to defeat.

Consuming strategies are a book of single-use tickets; reusing strategies are a season pass. Gatling sells only those two products — there is no ticket you borrow, use, and hand back for the next person.

saying these in an interview costs you the question

  • Choosing a strategy per scenario rather than per data file
  • Assuming random spreads evenly across keys in a finite run
  • Expecting a strategy that reuses records yet keeps them unique
  • Redeclaring the same file twice and losing the uniqueness guarantee