skip to content

Six JMeter engines stream samples and the controller's heap keeps filling. Which mode change helps most?

level: seniorimportance: must knowfreq 49%

answer

  1. Two dimensions, and they need different levers
  2. One mode reduces how many, not how big
  3. Spooling to disk helps the wrong node
  4. Aggregates need a count column to mean anything

basics

~20 s

Move to mode=Statistical. It is the only value that cuts the number of results reaching the controller rather than their size, replacing each label's samples with one aggregate. Stripping and batching only shrink or group what still arrives.

solid answer

~50 s

Confirm first that stripping is on: `mode` defaults to `StrippedBatch`, so if someone set `mode=Batch` to get response bodies, putting that back reclaims most of the heap immediately. If bodies are already stripped and the controller is still filling, the problem is the **count** of `SampleResult` objects, and only `mode=Statistical` reduces that — `StatisticalSampleSender` collapses each batch into one `StatisticalSampleResult` per sample label and thread group. Raising `num_sample_threshold` does not help: the same objects arrive, just in fewer, larger `processBatch` calls. `DiskStore` and `StrippedDiskStore` relieve the **engines**, not the controller; every sample still crosses at `testEnded`, one at a time, in a burst. Statistical costs you per-sample detail, so decide what you intend to report before you choose it. Which listeners you leave in the plan is a separate cost, owned by the plan-cost topic.

code

properties · 13 lines
properties
# controller's jmeter.properties
mode=Statistical

# aggregate on sample label + thread group (the default)
key_on_threadname=false

# the same two thresholds Batch uses, per aggregation window
num_sample_threshold=100
time_threshold=60000

# WITHOUT this, the JTL has no SampleCount/ErrorCount columns and each
# aggregate row's elapsed value looks like one enormously slow sample
jmeter.save.saveservice.sample_count=true

go deeper

for a junior

Recall that the mode property governs what engines send home, and that the default already removes response bodies. Recognising heap pressure as a sending-mode question is the step that matters here.

for a middle

Explain the difference between shrinking each result and reducing how many arrive, and say which mode value does which. Know that batching changes call granularity, not total volume.

for a senior

Work the order of moves: verify what is actually configured, raise heap before discarding detail, and only then trade per-sample rows for aggregates. Know that DiskStore relieves the engine, not the controller.

for a principal

Own the trade between fleet size and reportable detail. Decide in advance which runs may use Statistical, since the per-sample data it discards cannot be recovered and the decision silently constrains what the report can claim.

## What is actually arriving Six engines with a `Remoteable` listener in the plan each hold a `RemoteListenerWrapper` whose sender pushes `SampleEvent` objects back to the controller over RMI. The controller's listeners hold each result until they have written it. Two dimensions drive the heap: - **size per result** — dominated by `responseData`; - **number of results** — one per sample, per engine, for the whole run. `mode` is the lever that acts on both, and it acts on them differently. That distinction is the whole answer. ## What each change does to the controller | Change | Size per result | Number of results | Relieves | |---|---|---|---| | `mode=StrippedBatch` (default) | bodies removed | unchanged | controller and network | | `mode=Statistical` | small aggregates | **greatly reduced** | controller | | raise `num_sample_threshold` | unchanged | unchanged | per-call overhead only | | `mode=Asynch` | unchanged | unchanged | the engine's sampler threads | | `mode=DiskStore` / `StrippedDiskStore` | unchanged / stripped | unchanged | the **engine**, not the controller | Two rows in that table are the ones people get wrong. Raising the batch threshold *feels* like it should help the controller, but `BatchSampleSender` sends the same events in fewer calls — and a bigger clone list per call, so peak transient allocation actually rises. And `DiskStore` spools to a temp file **on the engine**, then in `testEnded` reads the file back and calls `sampleOccurred` for every event, one at a time. It converts a steady stream into an end-of-run flood, which is the worst possible shape for a controller that is already short of heap. ## Statistical is the only volume reduction `StatisticalSampleSender` keys each incoming result on `sampleLabel` plus thread group (or thread name, if `key_on_threadname=true`) and folds it into a `StatisticalSampleResult`. What the controller receives per key, per batch, is one object carrying: - summed elapsed time, returned by an overridden `getTime()`; - summed latency and connect time; - summed `bytes` and `sentBytes`; - a sample count and an error count; - earliest start time and latest end time. Everything that varies between samples — individual timings, response codes, messages, URLs, per-sample timestamps — is gone. Whether that loss is acceptable for the statistic you intend to report is a performance-testing fundamentals question and not a JMeter one; what JMeter guarantees is only that the individual rows no longer exist anywhere downstream. ## The Statistical trap you must configure around `jmeter.save.saveservice.sample_count` defaults to **`false`**, so a CSV JTL written under `Statistical` mode has no `SampleCount` or `ErrorCount` column — and because `getTime()` returns the **sum**, every row looks like one absurdly slow sample. Turn the property on before you run: ``` mode=Statistical key_on_threadname=false jmeter.save.saveservice.sample_count=true ``` Without it the run completes and the numbers are nonsense. ## A workable order of moves 1. **Check `mode` is not `Batch` or `Standard`.** Somebody wanting response bodies is the most common cause; restoring `StrippedBatch` is free. 2. **Check `sample_sender_strip_also_on_error`.** If it was set to `false` to keep failure bodies and the run is failing heavily, every failed body is crossing the wire. 3. **Give the controller more heap** before you give up detail — a mode change is irreversible for that run's data, an `-Xmx` change is not. 4. **Switch to `Statistical`**, with `sample_count` saving enabled, when the result count itself is the ceiling. 5. **Do not reach for `DiskStore`** hoping to help the controller; use it only when an *engine* is the one running out of memory.

  • Would raising num_sample_threshold to 5000 relieve the controller?
    No. `BatchSampleSender` still sends every `SampleEvent`; it just packs more into each `processBatch` call. The controller receives the same objects, the engine holds a longer list between sends, and the per-call clone gets bigger, so peak transient allocation on both sides rises.
  • Why does DiskStore not help a controller that is out of heap?
    `DiskStoreSampleSender` serialises events to a temp file on the **engine**. At `testEnded` it reads the file back and calls `sampleOccurred` for each one, so the controller receives everything anyway — all of it at the end of the run, as a burst rather than a stream.
  • What must you enable before a Statistical run to get a readable JTL?
    `jmeter.save.saveservice.sample_count=true` on the controller. It defaults to false, and without the `SampleCount` and `ErrorCount` columns you cannot tell how many samples each aggregate row stands for, while its elapsed value is their sum rather than an average.
  • How do you tell whether stripping is even active on this run?
    Each sender logs its own name at start-up. `DataStrippingSampleSender` writes "Using DataStrippingSampleSender for this run" with the resolved `stripAlsoOnError` value into the engine's `jmeter-server.log`, so the log tells you what was really selected rather than what you meant to select.

saying these in an interview costs you the question

  • Raises num_sample_threshold expecting the controller to receive less
  • Reaches for DiskStore to relieve controller memory
  • Thinks Asynch mode reduces what the controller receives
  • Switches to Statistical without enabling sample_count saving
  • Assumes stripping is off because response bodies were never requested
  • Treats a mode change as reversible for data already discarded