skip to content

In k6, which aggregations may a Trend threshold use, and why is `count<100` rejected on one?

level: middleimportance: nice to knowfreq 37%

answer

  1. five aggregations, no more
  2. trends are the only type keeping samples
  3. percentile argument runs 0 to 100
  4. med is the fiftieth percentile
  5. tallies are a counter's job

basics

~20 s

A k6 Trend threshold accepts exactly avg, min, max, med and p(N) with N from 0 to 100. count is not among them, because counting samples belongs to the Counter type, so count<100 on a Trend is rejected as invalid configuration.

solid answer

~40 s

The Trend row of k6's aggregation table is `avg`, `min`, `max`, `med` and `p(N)` — five entries and no more. `N` may be any number from 0 to 100 inclusive, including fractional values such as `p(99.9)`, and `med` is simply k6's name for the 50th percentile, so `med<200` and `p(50)<200` evaluate identically. `count` is deliberately absent: sample counting is the Counter type's job, and a Trend's supported list does not include it, so `http_req_duration: ['count<100']` fails validation and the test does not start. The confusion is understandable, because k6's end-of-test summary *can* display a sample count for a Trend — but the summary's statistics vocabulary and the threshold grammar are separate sets, and only the latter governs what you may write in `thresholds`.

code

javascript · 16 lines
javascript
export const options = {
  thresholds: {
    http_req_duration: [
      'avg<200',
      'min<50',
      'max<3000',
      'med<150',      // identical to 'p(50)<150'
      'p(95)<400',
      'p(99.9)<2000', // fractional percentiles are fine
    ],

    // Rejected before the run starts:
    // http_req_duration: ['count<100'],   // count is a Counter method
    // http_req_duration: ['p(101)<400'],  // percentile out of range
  },
};

go deeper

for a junior

Memorise the five: avg, min, max, med and p(N). If a threshold on a k6 duration metric uses anything else, the test will not start.

for a middle

Explain that the percentile argument is validated during parsing, that med is an alias for the 50th percentile, and that count belongs to the Counter type.

for a senior

Point out that a Trend threshold on an unsampled metric still evaluates against zero, so a ceiling rule can report green on a run that produced no traffic at all.

for a principal

The judgement to surface is which statistics a team standardises on across scripts, given that k6 will compute any percentile you name but only the ones you name become visible in the output.

## The Trend row in full A Trend is the only k6 metric type that retains individual samples, which is why it is the only one that can answer distribution questions. Its threshold vocabulary is exactly five entries: | Aggregation | Meaning | |---|---| | `avg` | arithmetic mean of every recorded sample | | `min` | smallest recorded sample | | `max` | largest recorded sample | | `med` | median, computed as the 50th percentile | | `p(N)` | the Nth percentile, `N` between 0 and 100 | Every built-in timing metric is a Trend — `http_req_duration`, `http_req_waiting`, `iteration_duration`, `group_duration`, `grpc_req_duration` — and so is any metric you declare with `new Trend(...)`. ## The percentile argument `p(N)` is the only parametric aggregation in the whole grammar, and its argument is validated while the expression is parsed, before k6 ever looks at which metric the rule applies to. - `N` is parsed as a floating-point number, so `p(99.9)` and `p(99.99)` are both fine. - The accepted interval is closed at both ends: `p(0)` and `p(100)` are legal. - Anything outside it fails to parse — `p(101)`, `p(-1)`, `p(NaN)` and `p(+Inf)` all stop the test as an invalid configuration. - Malformed shapes fail too: `p()`, `p(foo)` and an unclosed `p(99` are all parse errors. You are not restricted to the percentiles the summary happens to print by default. k6 computes each percentile that appears in a threshold on demand from the retained samples, and it will also add that percentile to the metric's displayed statistics in the end-of-test summary if it was not already there — so declaring `p(99.9)<2000` makes `p(99.9)` visible in the output as well as enforced. ## Why `med` and `p(50)` are the same thing k6 resolves `med` by asking the trend sink for the 50th percentile, exactly as it would for `p(50)`. The two expressions therefore compare identical numbers. They are distinct sink keys internally, so declaring both is harmless but redundant; `med` exists purely as a readable alias. ## Why `count` is rejected The instinct to write `http_req_duration: ['count<100']` comes from somewhere real: a Trend does know how many samples it holds, and k6's end-of-test summary can display that number. But the two vocabularies are separate sets, and the threshold grammar takes its Trend list from the metric type's supported-aggregation list, which contains no `count`. What you get is a validation failure naming the unsupported method and listing the five that are supported, and no test run. If you actually want to bound the number of samples, use a Counter — `http_reqs` for requests, `iterations` for completed iterations — because counting volume is precisely what that type is for. ## What each Trend aggregation is computed from A Trend sink retains every value handed to it, along with a running sum, a running minimum and a running maximum. That storage decides how each aggregation is answered: - `min` and `max` come straight from the running extremes, so they cost nothing to evaluate. - `avg` is the running sum divided by the number of samples. - `med` and every `p(N)` require the retained samples to be sorted, and k6 interpolates linearly between the two neighbouring samples when the requested position falls between them. This is also why the Trend type is the expensive one: retaining samples is what buys you percentiles, and no other k6 metric type keeps enough state to offer them. A Counter has a single running total, a Gauge remembers only its latest sample, and a Rate holds two integers. Asking any of those for a percentile is not a limitation k6 chose to impose — there is nothing in the sink to compute it from. One consequence is worth planning for. Because a Trend threshold on a metric that never received a sample still resolves `min`, `max`, `avg` and `med` to zero rather than being skipped, a ceiling such as `p(95)<400` reports a pass on a run that measured nothing at all. If the point of the rule is to prove the system was exercised, pair it with a Counter floor such as `http_reqs: ['count>0']`. ## A short checklist 1. Confirm the metric is a Trend; only then are percentiles available. 2. Choose from `avg`, `min`, `max`, `med`, `p(N)` and nothing else. 3. Keep `N` inside 0 to 100, and remember fractional percentiles are allowed. 4. Reach for a Counter, not a Trend, when the quantity you want is a tally. One footnote on style rather than legality: `min` and `max` are fully supported, but k6's own documentation marks them as not recommended for thresholds. They will run; the caution is about what they measure, not about whether k6 accepts them.

  • Can a k6 threshold use a percentile that the end-of-test summary does not normally show?
    Yes. k6 computes a percentile on demand from the trend's retained samples for every `p(N)` that appears in a threshold, so `p(99.9)<2000` is enforced regardless of which statistics the summary prints by default. k6 also folds that percentile into the metric's displayed values, so the enforced figure appears in the output next to the rule.
  • How does k6 treat a Trend threshold on a metric that received no samples?
    The trend sink still resolves `min`, `max`, `avg` and `med` to zero, so the rule is evaluated against zero rather than skipped. A ceiling such as `p(95)<200` therefore reports green on a run that measured nothing, while a floor such as `avg>100` breaches. During the run k6 skips metrics whose sink is still empty, so the outcome only appears on the final evaluation.

saying these in an interview costs you the question

  • Writing count on a Trend because the summary shows a sample count
  • Believing percentiles are limited to the ones printed by default
  • Thinking med and p(50) compute different statistics
  • Assuming p(101) is clamped to 100 rather than refused
  • Expecting rate or value to work on a duration metric