Do k6's p(95) latency and http_req_failed rate thresholds pass on a run that made zero requests?
answer
- the final pass judges everything
- empty sinks skipped only in-run
- missing value counts as passing
- an empty trend reports zero
- add a lower bound on a count
basics
~20 sBoth pass. k6's final evaluation judges every declared threshold, and a metric with no samples offers the comparison either nothing at all, which counts as a pass, or a zero that beats any upper bound.
solid answer
~40 sThey pass, and that is the failure mode worth knowing. The final evaluation does not skip metrics that never received data. For `http_req_failed: ['rate<0.01']` the sink has no observations, so it hands the comparison nothing at all and k6 treats the rule as satisfied. For `http_req_duration: ['p(95)<500']` an empty trend reports zero for every aggregation, and `0 < 500` is true. An upper-bound rule therefore cannot distinguish "fast" from "never ran". The counterpart is a lower-bound rule on a count: a counter always reports its total, so `http_reqs: ['count>0']` genuinely fails on an empty run, which is why one is worth adding.
code
javascript · 9 linesexport const options = {
thresholds: {
// both of these pass on a run that made no requests at all
http_req_duration: ['p(95)<500'],
http_req_failed: ['rate<0.01'],
// this one does not: a counter always reports its total
http_reqs: ['count>0'],
},
};go deeper
Recall that a k6 threshold judges an aggregate at the end of the run, and that an aggregate over no data does not mean the same thing as a good result.
Explain the difference between the in-run evaluation, which skips metrics with no samples, and the final one, which does not - and what an empty sink offers the comparison in each case.
Recognise the shape in the wild: a suite whose latency gate has been green for weeks because the requests stopped being made. Add a rule that cannot pass on an empty set and say why.
Treat a rule that absence can satisfy as a class of defect rather than one bug, and decide whether every k6 script in the suite carries a count-based guard by convention and where that rule is kept.
## Two evaluation passes, not one k6 evaluates thresholds twice in different modes, and the difference is the whole answer: - **During the run**, on a repeating tick, k6 evaluates thresholds but **skips any metric whose sink is still empty**. No data, no evaluation, no failure. - **At the end of the run**, it evaluates once more to fix the verdict - and this pass does **not** skip empty sinks. Every metric that has a threshold declared on it is evaluated, whether or not it ever saw a sample. So a metric that never got data is silently ignored while the test runs, then judged at the end. What it is judged on depends entirely on what an empty sink can offer the comparison. ## What an empty sink hands the expression Before comparing, k6 collects the metric's aggregations into a small lookup and then looks up the one your expression named. Two outcomes are possible, and they behave differently: - **The name is missing from the lookup.** k6 treats this as "no samples yet, nothing to judge" and the expression is recorded as **passing**. - **The name is present with a zero value.** The comparison runs normally against `0`. Which of the two you get is a property of the metric's own shape: | threshold on an empty run | what the lookup holds | result | |---|---|---| | `http_req_failed: ['rate<0.01']` | nothing - no observations were recorded | **passes** | | `http_req_duration: ['p(95)<500']` | `0` for every aggregation of an empty trend | **passes** (`0 < 500`) | | `http_reqs: ['count>100']` | `0` - a counter always reports its total | **fails** | | `http_reqs: ['count>0']` | `0` | **fails** | The two rules from a normal latency-and-errors budget land in the two passing rows. Nothing in the summary flags this: both appear with a green tick, the run exits successfully, and a pipeline gate downstream sees a healthy result. ## Why this bites in practice The run does not have to be literally empty for this to matter. Any of these produce the same shape: 1. **A script that throws in its setup or in every iteration** before the first request goes out. The HTTP metrics stay empty; the latency and error-rate rules pass vacuously. 2. **A wrong base URL or an unresolvable host** in an environment where every request fails at a layer that never records a duration sample. 3. **A scenario that never starts** - a start time past the end of the test, or an executor configured so that no iteration is ever scheduled. 4. **A tagged rule whose selector matches nothing** because the tag was renamed. That sub-metric collects no samples, and its upper-bound rule passes on every run from then on. In every case the threshold you wrote to protect latency is not protecting anything - it is asserting a property of an empty set, and an upper bound on an empty set is trivially true. ## The shape of the fix The asymmetry in the table above is also the way out. **A rule that fails when a value is zero cannot pass vacuously**, so pair every upper-bound rule with a lower bound on something that only traffic can move: ```javascript export const options = { thresholds: { http_req_duration: ['p(95)<500'], http_req_failed: ['rate<0.01'], http_reqs: ['count>0'], }, }; ``` `http_reqs` is a counter, and a counter's total is always available to the comparison - zero included - so `count>0` is false on an empty run and the whole run is marked failed. The same trick works with a custom counter you increment yourself when a code path you care about actually executes; put a `count>0` rule on it and a silently skipped path stops being invisible. ## Two things this is not It is **not** a warm-up problem. During the run, skipping empty sinks is deliberate and helpful - it stops a threshold from failing on the two or three samples that happen to exist at the first evaluation. The vacuous pass only appears at the final evaluation, where skipping would be worse still, because a metric with a threshold and no data would then vanish from the verdict entirely. It is also **not** something the parser can catch for you. `'p(95)<500'` is a perfectly valid expression; k6 has no way to know that you intended it to imply "and there was traffic". The implication has to be written down as its own rule.
- Why does k6 skip empty sinks during the run but not at the end?During the run the skip protects a threshold from being judged on the two or three samples that exist at the first evaluation. At the end there is nothing left to wait for, and skipping would drop a declared threshold out of the verdict altogether.
- Does a threshold on a sub-metric that matches nothing behave the same way in k6?Yes, and it is the version of this you are most likely to ship. The sub-metric is created from the threshold declaration itself, so it always exists; it simply never collects samples, and its upper-bound rule passes on every run.
- Would a rate threshold written as 'rate==0' catch an empty k6 run?No. With no observations there is no rate value for the comparison to read at all, so k6 records the expression as passing regardless of which operator you chose. Only a metric that still reports a number when empty - a counter - can be made to fail.
saying these in an interview costs you the question
- assumes a green threshold proves the run generated traffic
- thinks a metric with no samples fails its threshold
- believes the end-of-run pass also skips empty metrics
- relies only on upper-bound rules to gate a run
- expects an equality operator to catch the empty case