skip to content

A JMeter CSV Data Set Config feeds unique logins across six engines - why does every engine issue the same rows?

level: seniorimportance: must knowfreq 62%

answer

  1. One file service per JVM
  2. The alias never leaves the engine
  3. Six readers, six line ones
  4. Split the data before the run, not during

basics

~20 s

Every engine holds its own copy of the file and its own read position, so all six start at line one and walk the same rows in step. Sharing mode scopes threads inside one JVM and never reaches across engines.

solid answer

~50 s

JMeter's file service is a singleton **per JVM**. It opens a file once per alias and keeps one read position for it, and the CSV Data Set Config's **Sharing mode** only decides which alias a thread uses: the bare filename for *All threads*, the filename plus the thread group's identity for *Current thread group*, plus the thread's identity for *Current thread*, or plus a suffix you type. Every one of those aliases lives inside a single engine's JVM. Because the controller replicates the plan, six engines each reserve the same file independently, each from line one, and 1,000 threads on each engine consume rows 1..1000 six times over. The data is duplicated, not exhausted six times faster. The fix is to make the rows disjoint before the run: split the file into per-engine slices, or give the filename a per-engine name.

code

xml · 8 lines
xml
<CSVDataSet guiclass="TestBeanGUI" testclass="CSVDataSet" testname="CSV Data Set Config" enabled="true">
  <stringProp name="filename">data/logins-${__machineName}.csv</stringProp>
  <stringProp name="variableNames">user,password</stringProp>
  <stringProp name="delimiter">,</stringProp>
  <stringProp name="shareMode">shareMode.all</stringProp>
  <boolProp name="recycle">false</boolProp>
  <boolProp name="stopThread">true</boolProp>
</CSVDataSet>

go deeper

for a junior

Know that each engine reads its own copy of the file from the top. Copying the same data file to six servers gives six identical sequences, not one shared queue of rows.

for a middle

Explain the mechanism: the file service is a per-JVM singleton keyed by an alias, and Sharing mode only picks the alias. No alias crosses an engine boundary, so no position is shared.

for a senior

Show that you make the rows disjoint at deploy time and that you verify it - per-engine checksums before, distinct-value counts after. Talk about what the duplicates did to the target, not just to the data.

for a principal

Own the data pipeline. Decide how test data is generated, partitioned and refreshed per engine, and make the partition survive a change in the number of engines rather than being re-cut by hand each run.

## Why the rows repeat Two facts collide. First, the controller sends every engine the same plan, so all six run the same CSV Data Set Config element. Second, the file service that element uses is a **static singleton inside each JVM** - it reserves a file under an alias, opens one reader for it, and hands out the next line to whoever asks. Neither fact knows the other exists. Six engines therefore open six readers on six local copies of the file and every one of them starts at line one. So a 6,000-row file feeding 1,000 threads on each of six engines does not run out after 6,000 rows; each engine consumes rows 1 to 1000 and stops there, and every login is used six times concurrently. If those rows are supposed to be unique - one account per virtual user, one order id per checkout - the run tests a *different workload* from the one you designed, and often a much heavier one, because six sessions are now contending for the same account row. ## Sharing mode stops at the JVM boundary | Sharing mode | Alias the element uses | Scope in one engine | Scope across six engines | |---|---|---|---| | All threads (default) | the filename | one position for the whole engine | six independent positions | | Current thread group | filename + the group's identity | one per thread group | six per thread group | | Current thread | filename + the thread's identity | one per thread | six per thread | | A suffix you type | filename + your suffix | one per suffix | six per suffix | Read the last column twice. **There is no sharing mode that spans engines**, because the mechanism is an alias in a per-JVM map and nothing replicates that map. Picking *All threads* buys you one position per engine, which is the tightest scope distribution allows - and it is still six positions. ## Two ways to make the rows disjoint 1. **Slice the file and place one slice per engine.** Cut the source data into as many disjoint files as you have engines, name each one identically, and deploy slice *n* to engine *n* at the same relative path. The plan is unchanged; the data differs per box. The manual recommends exactly this: *"use different content in any datafiles used by the test (e.g. if each server must use unique ids, divide these between the data files)".* 2. **Give the filename a per-engine name.** Because the plan body is resolved on the engine that runs it, a function that reports the local machine resolves differently on each one. `logins-${__machineName}.csv` in the Filename field makes each engine open its own file, and the files can hold disjoint rows. This costs you a naming convention that has to survive a host being replaced. Both approaches move the split out of JMeter and into your deployment, because JMeter has no run-time step that could do it. ## What the duplicates do to your numbers - **Server-side contention you did not ask for.** Six sessions on one account row serialise on locks the real workload would never hit. - **False cache hits.** The same six ids are requested over and over, so hit rates and response times flatter the system. - **Uniqueness violations.** Any row that creates something with a unique key fails on five of six engines, and the failures may not be asserted. - **A row budget that looks right and is not.** The file has 6,000 rows and the run wanted 6,000 users, so nothing looks short. ## Checks before the run 1. Multiply threads by engines, then confirm the *per-engine* file holds at least that many usable rows. 2. Grep the result file for a value that must be unique and count distinct occurrences; six times the expected repetition is the signature. 3. Verify the deployed slices differ - compare a checksum of the data file on each engine. Identical checksums on a plan that needs disjoint data is the defect. 4. Decide in advance what should happen at the end of a slice, and make that behaviour explicit on the element rather than leaving it to the default.

  • Does setting Sharing mode to All threads make six JMeter engines share one read position?
    No. The mode selects an alias inside one JVM's file service, and that map is not replicated. *All threads* gives one position per engine, which is the tightest scope distribution allows - six engines still mean six positions, all starting at line one.
  • Your six per-engine CSV slices are unequal because one engine was added late. What breaks first?
    The engine with the short slice reaches the end of its file first and behaves according to the element's end-of-file settings - recycling into duplicate rows or stopping its threads early. Either way that engine stops applying the load the other five still are, and the run's applied total drops mid-flight.
  • How would you prove after a run that the engines used disjoint rows?
    Save the identifying variable with each sample and count distinct values against the total. If the count is roughly the row total divided by the engine count, the engines overlapped. Comparing a checksum of the deployed data file per engine catches the same defect before the run.

saying these in an interview costs you the question

  • Thinks Sharing mode All threads spans the remote engines.
  • Assumes a 6,000-row file lasts a 6,000-thread distributed run.
  • Expects the controller to hand out rows to the engines.
  • Copies one identical data file to every engine and calls the rows unique.
  • Blames the target's locking for contention the duplicated rows caused.