A JMeter gate counting false rows in the JTL passes when the run produced no rows. How do you harden it?
answer
- The numerator was fine; look at the denominator
- Absence of evidence is passing this rule
- An empty results file still has one line
- Check the artefact before checking the ratio
basics
~20 sAdd a floor as well as a ceiling. Check that the results file exists, that it holds more than its header line, and that the sample count is near what the plan should produce, before computing any ratio.
solid answer
~40 sA rule shaped as *fail if false rows exceed one percent* divides by a total that can legitimately be zero, and three JMeter behaviours get you there. A test-tree compile error returns from `StandardJMeterEngine.run()` before any listener is told the test started, so `ResultCollector` never opens the file and `-l` leaves nothing at all. A run that started but sampled nothing still leaves a one-line file, because `jmeter.save.saveservice.print_field_names` defaults to `true`. And `jmeter.save.saveservice.autoflush` defaults to `false`, so a job killed on a timeout leaves a truncated JTL whose ratio flatters the run. Order the checks and let each fail on its own: file present and non-empty, data rows above an expected floor, then the error ratio — under `set -euo pipefail`, so the check's own failure is never read as a pass.
code
bash · 17 linesset -euo pipefail
JTL=results.jtl
MIN_SAMPLES=1000
MAX_ERROR_PCT=1
[ -s "$JTL" ] || { echo "FAIL: JMeter never wrote $JTL"; exit 1; }
read -r header < "$JTL"
col=$(awk -v h="$header" 'BEGIN { n = split(h, f, ","); for (i = 1; i <= n; i++) if (f[i] == "success") print i }')
[ -n "$col" ] || { echo "FAIL: no success column in $JTL"; exit 1; }
read -r total errors < <(awk -F, -v c="$col" 'NR > 1 { t++; if ($c == "false") e++ } END { print t + 0, e + 0 }' "$JTL")
[ "$total" -ge "$MIN_SAMPLES" ] || { echo "FAIL: only $total samples, expected at least $MIN_SAMPLES"; exit 1; }
[ $(( errors * 100 )) -le $(( total * MAX_ERROR_PCT )) ] || { echo "FAIL: $errors of $total samples failed"; exit 1; }
echo "PASS: $errors of $total samples failed"go deeper
Understand that an empty results file is not the same as a clean run, and that a check has to look at how many samples were recorded before it looks at how many failed.
Explain the three ways a JTL ends up with no usable rows: no file at all after a compile error, a header-only file, and a truncated file from an unflushed buffer.
Design the check as ordered assertions that each fail with their own message, resolve the column by header name, and fail closed when the check itself cannot read what it needs.
Set the estate-wide rule that a load stage must prove it ran before it is allowed to prove it passed, and decide who owns the sample floor for each plan as the plan changes.
## The shape of the bug The rule is right about the numerator and silent about the denominator. `errors / total <= 0.01` with `total == 0` either divides by zero or, in the shell, evaluates `0 -le 0` and passes. Every false-green story on this leaf is a variant of it: the run that failed every sample at least produced rows to count, but the run that produced **no** rows produced no evidence either, and an absence looks exactly like a clean result to a rule that only counts errors. ## Three JMeter behaviours that reach zero rows 1. **The results file may not exist.** `StandardJMeterEngine.run()` compiles the plan with `PreCompiler` inside a `try`; on a `RuntimeException` it logs `Error occurred compiling the tree:` and returns *before* notifying test listeners. `ResultCollector` opens its file in `testStarted`, which is therefore never called. The process still exits `0` and `-l` wrote nothing. 2. **A file can exist with no samples in it.** When the run does start, `ResultCollector.getFileWriter` writes the CSV header immediately, and `jmeter.save.saveservice.print_field_names` defaults to `true`. So an empty run leaves a file of exactly one line. `[ -s "$JTL" ]` passes on it, and `wc -l` returns 1, not 0. 3. **A file can be truncated.** The writer wraps a `BufferedOutputStream` and `jmeter.save.saveservice.autoflush` defaults to `false`. Kill the JVM — a job timeout, an OOM, a runner reclaim — and whatever was in the buffer is lost. The surviving rows are a *sample* of the run, and if the failures clustered late they are the rows you lost. ## The order the checks have to run in - **Artefact exists and is non-trivial.** Fail if the file is missing or empty; fail separately if it has no data rows, and say which, because the two mean different things to whoever reads the log. - **Sample count meets a floor.** The floor comes from the plan: threads times loops, or the rate times the duration. A run that produced a tenth of the expected samples did not pass, whatever its error ratio says. - **Then the error ratio.** Only now is the denominator meaningful. - **Resolve the column by name.** The JTL layout is property-driven, so `success` is only the eighth field under a default `jmeter.save.saveservice.*` configuration. Read the header line and find the index; a hard-coded `$8` breaks silently the day someone adds a column. - **Fail closed on the checker's own errors.** `set -euo pipefail`, and treat "no `success` column found" as a failure rather than as zero errors. A gate whose input went missing must go red. ## The parsing caveat `CSVSaveService` quotes any field containing the delimiter, so a URL with a comma in a query string or an assertion message containing one produces a quoted field that a plain `awk -F,` will split in the wrong place. That is tolerable for a smoke check on a plan you control, and not tolerable as a shared gate — use a real CSV reader, or set the delimiter to something the data cannot contain, once the rule is guarding more than one plan. ## What is still outside the JTL's reach A truncated file and a job that hit its wall-clock limit look similar from inside the results file, so keep the job's own timeout as a separate failure signal rather than inferring it from row counts. And decide explicitly what a sub-sample means to your ratio: embedded resources and transaction children can appear as rows, so filter by `label` if the rule is meant to be about named transactions rather than about every HTTP call the plan made.
- Where should the expected minimum sample count come from?From the plan's own arithmetic rather than from last night's number. Threads times loops, or the target rate times the scheduled duration, gives a figure you can defend; a floor copied from the previous run drifts downward every time the run degrades, which is the failure it is supposed to catch.
- Why can a JTL be shorter than the run that produced it?Because jmeter.save.saveservice.autoflush defaults to false and the result collector writes through a BufferedOutputStream. If the JVM is killed rather than shut down, the buffered rows are lost, so a timed-out job leaves a partial file whose error ratio can look better than the run really was.
- Should the gate parse the JTL with awk?For one plan you own, a field-split smoke check is fine. As a shared rule it is not: CSVSaveService quotes any field containing the delimiter, so a comma in a URL or an assertion message shifts the columns. Use a real CSV reader, or pick a delimiter the data cannot contain.
saying these in an interview costs you the question
- Computes an error ratio without checking the sample count
- Treats a missing results file as zero errors
- Hard-codes the success column as field eight
- Runs the check without set -e, so its own errors pass
- Assumes a JTL is always complete when the job ends