skip to content

Over which samples does a JMeter Backend Listener compute the percentiles it streams?

level: seniorimportance: should knowfreq 42%

answer

  1. Not the interval you would assume
  2. A count of samples, not a duration
  3. Two properties, two different windows
  4. One hundred is the shipped size

basics

~10 s

Over a fixed-size sliding window of the most recent response times - 100 values by default, set by backend_metrics_window - and never over the interval just elapsed. That window is not cleared between flushes.

solid answer

~40 s

`SamplerMetric` holds the *all*-status percentiles in a dedicated `pctResponseStats` sized by `backend_metrics_window`, default `100`, which no window mode ever clears; the source comment is blunt - *Timeboxed percentiles don't makes sense*. So `a.pct99` (Graphite) and the `statut=all` percentiles (InfluxDB) describe roughly the last hundred samples the listener saw, not the last five seconds. The `ok` and `ko` percentiles are different: `getOkPercentile`/`getKoPercentile` read `okResponsesStats`/`koResponsesStats`, the same statistics behind min, max and average, and those follow `backend_metrics_window_mode` (default `fixed`) - in `fixed` mode a 100-value sliding window that is never reset, in `timed` mode cleared every interval and holding up to `backend_metrics_large_window` values, default `5000`. Only the counters - `count`, `countError`, `hit`, `sb`, `rb` - are per-interval in both modes. On a busy run, a hundred samples can be well under a second of traffic.

code

properties · 14 lines
properties
# bin/jmeter.properties - Backend Listener statistics windows (shipped commented out)

# fixed = fixed-size sliding window, timed = time boxed
#backend_metrics_window_mode=fixed

# sliding window size for Percentiles, Min, Max
#backend_metrics_window=100

# sliding window size for Percentiles, Min, Max when backend_metrics_window_mode is timed
# Setting this value too high can lead to OOM
#backend_metrics_large_window=5000

# how often the InfluxDB client flushes, in seconds
#backend_influxdb.send_interval=5

go deeper

for a junior

Know that the streamed percentile is an estimate over recent samples and that a JMeter property controls how many are kept. Do not read it as the run's overall percentile.

for a middle

Explain the split: counts reset every flush, response-time statistics live in a sliding window. Name backend_metrics_window and its default of 100, and what the fixed and timed modes each change.

for a senior

On a long run, be able to say why the live graph stayed flat while the end-of-run figure was far worse, and what you would change - window size, window mode - before trusting the live line.

for a principal

Decide what the live graph is allowed to be used for. If a soak's exit criteria are read off a dashboard, window sizing is part of the test design and not a display preference.

## Two windows, not one interval `SamplerMetric` is the class behind every number a Backend Listener sends, and it keeps response times in more than one place: - **`pctResponseStats`** - the source of the *all*-status percentiles only (`a.pctNN` in Graphite, the `statut=all` percentiles in InfluxDB). It is a `DescriptiveStatistics` sized by `backend_metrics_window`, **default 100**, and it is **never cleared**, in either window mode. The comment above it in the source says only: *Timeboxed percentiles don't makes sense*. - **`okResponsesStats` / `koResponsesStats` / `allResponsesStats`** - the source of `ok.min`, `ok.max`, `ok.avg`, their `ko` and `a` counterparts, **and the `ok.pctNN` / `ko.pctNN` percentiles**. What happens to these depends on `backend_metrics_window_mode`. - **Plain counters** - `successes`, `failures`, `hits`, `sentBytes`, `receivedBytes` and the error map. These are reset on every flush, in both modes. So a single flush mixes two time bases: counts that describe the interval just ended, and response times that describe a rolling window of recent samples. ## What each window mode changes | | `fixed` (the default) | `timed` | |---|---|---| | ok/ko/all min, max, avg | sliding window of `backend_metrics_window` values, never cleared | cleared every flush, holding up to `backend_metrics_large_window` values (default 5000) | | all-status percentiles (`a.pctNN`, `statut=all`) | sliding window of `backend_metrics_window` values, never cleared | identical - the mode does not touch it | | ok / ko percentiles | sliding window of `backend_metrics_window` values, never cleared | cleared every flush, holding up to `backend_metrics_large_window` values | | counts, bytes, hits, errors | reset every flush | reset every flush | The rows that surprise people are the two percentile ones. Switching to `timed` does **not** make the all-status percentile per-interval - that one is read from a statistic no mode clears. It does time-box min, max, average **and the `ok`/`ko` percentiles**, which share those statistics. ## The soak that looks calm Picture a twelve-hour soak whose progress you are following on a dashboard fed by the InfluxDB client. Response times drift upward over the night. At 2 a.m. the plotted `pct99` still looks respectable, and the end-of-run report says something much worse. Nothing is broken. At, say, 200 samples per second, one hundred samples is half a second of traffic. The plotted percentile is an estimate over the last fraction of a second the listener happened to see, redrawn every five seconds. It reacts fast to a sudden change and it says nothing at all about the twelve hours behind it. The counts on the same dashboard - `count`, `countError`, `hit` - *are* per-interval and *are* trustworthy as interval quantities. ## What to do about it 1. **Raise `backend_metrics_window`** if the graph is too twitchy at the sample rate you actually run. More values in the window means a steadier line and a longer memory. 2. **Do not raise it carelessly.** The property file's own note on the large window - *setting this value too high can lead to OOM* - is a warning about holding response times in memory, and the window is held per tracked label. 3. **Keep the authoritative distribution somewhere else.** The streamed percentile is for watching; the run's own results file is what a verdict should be computed from afterwards. 4. **Read the percentile row you configured.** `percentiles` is a semicolon-separated list - `90;95;99` for the Graphite client, `99;95;90` in the InfluxDB client's shipped arguments - and fractional values such as `12.5` are allowed. Each value becomes its own metric per label. Deciding what a percentile is worth as evidence, and how to compare two runs by it, belongs to performance-testing practice rather than to JMeter. What belongs here is the narrower fact: this number came from a bounded count of recent samples, not from a stretch of time.

  • Does setting backend_metrics_window_mode to timed make percentiles per-interval?
    Only for the `ok` and `ko` percentiles. In `timed` mode `okResponsesStats` and `koResponsesStats` - which back min, max, average *and* `ok.pctNN`/`ko.pctNN` - are cleared each interval and their window grows to `backend_metrics_large_window`, default `5000`. The all-status percentile (`a.pctNN`, `statut=all`) is read from `pctResponseStats`, a separate `backend_metrics_window`-sized statistic that neither mode clears, so that one stays a sliding estimate over recent samples whichever mode you pick.
  • Which percentiles does a JMeter Backend Listener send if you leave the row alone?
    The Graphite client's default `percentiles` is `90;95;99`; the InfluxDB client's shipped row is `99;95;90`. The list is semicolon-separated and a value may be fractional, such as `12.5`. Each value becomes its own metric for each label being tracked.

A car's instantaneous consumption readout averages the last stretch of road, not the last minute on the clock and not the whole journey. On a steady cruise it looks reassuring, while the trip computer's figure for the same drive can be far worse.

saying these in an interview costs you the question

  • Assumes the percentile covers the flush interval
  • Thinks the percentile spans the whole run so far
  • Believes counts and response times reset together
  • Expects the live percentile to match the final report
  • Thinks timed mode makes the all-status percentile per-interval