In k6, which threshold aggregations does each metric type allow, and what happens on a mismatch?
answer
- the type constrains the expression
- four types, four vocabularies
- counters keep no distribution
- parse first, then type-check
- invalid configuration, not a failed run
basics
~20 sIn k6 v2 a Counter takes count and rate, a Gauge takes value, a Rate takes rate, and a Trend takes avg, min, max, med and p(N). Any other pairing is rejected as invalid configuration before the test starts.
solid answer
~40 sk6 derives the legal left-hand side of a threshold expression from the metric's **type**, not from the metric's name. A **Counter** accepts `count` and `rate`; a **Gauge** accepts `value`; a **Rate** accepts `rate`; a **Trend** accepts `avg`, `min`, `max`, `med` and `p(N)` for any `N` from 0 to 100. k6 handles a threshold in two stages: it first parses the expression for shape, then validates the parsed aggregation against the type of the metric named by the key. If the type does not support the method, k6 reports an invalid threshold naming the offending expression, the metric, and the methods the type does support, and refuses to start the run — so `http_reqs: ['p(95)<200']` never reaches a virtual user, because `http_reqs` is a Counter and percentiles belong to Trends.
code
javascript · 18 linesimport { Trend, Rate, Counter, Gauge } from 'k6/metrics';
export const RTT = new Trend('RTT');
export const ContentOK = new Rate('ContentOK');
export const ContentSize = new Gauge('ContentSize');
export const Errors = new Counter('Errors');
export const options = {
thresholds: {
Errors: ['count<100', 'rate<5'], // Counter
ContentSize: ['value<4000'], // Gauge
ContentOK: ['rate>0.95'], // Rate
RTT: ['p(99)<300', 'avg<200', 'med<150'], // Trend
// Rejected before the run: http_reqs is a Counter.
// http_reqs: ['p(95)<200'],
},
};go deeper
Learn the four-row table by heart: Counter takes count and rate, Gauge takes value, Rate takes rate, Trend takes avg, min, max, med and p(N).
Explain that k6 parses the expression first and only then checks it against the metric's registered type, and that both failures stop the test before it starts.
Show that you read the error message: it names the offending expression, the metric's type, and the exact list of methods that type accepts, which is enough to repair the rule in one edit.
The angle to discuss is that binding aggregations to types trades expressiveness for a compile-time-style guarantee, and where a team wants percentiles it must choose the Trend type up front rather than convert later.
## The table is the whole rule Every metric in k6 — built-in or custom — has one of four types, and that type fixes which aggregation methods a threshold on it may use. There is no per-metric configuration and no override. | Metric type | Legal aggregations | What the value is | |---|---|---| | Counter | `count`, `rate` | the running sum, and that sum per second of elapsed run time | | Gauge | `value` | the most recently added sample | | Rate | `rate` | non-zero samples divided by total samples, in `[0, 1]` | | Trend | `avg`, `min`, `max`, `med`, `p(N)` | mean, extremes, median, and the Nth percentile | Mapping the built-ins onto it: `http_req_duration` is a **Trend**, so `p(95)<200` and `avg<150` are fine; `http_reqs` is a **Counter**, so `count>1000` and `rate>50` are fine; `checks` and `http_req_failed` are **Rates**, so only `rate` is available; `vus` and `vus_max` are **Gauges**, so only `value` is available. ## Two stages, two different failures k6 processes a `thresholds` entry in a fixed order: 1. **Parse.** The expression is split into an aggregation method, an operator and a numeric right-hand side. `foo&0` dies here, and so does `p(101)<200`, because a percentile argument must be a number from 0 to 100. 2. **Validate.** The metric named by the key is looked up in the registry, and the *parsed* aggregation is checked against that metric's type. Both stages run while the configuration is being assembled, before any executor starts, and both classify their failure as an invalid configuration rather than a test result. That distinction matters: a mismatch is not a failed threshold, so it does not produce the usual breached-threshold outcome. It produces a configuration error and no test run at all. ## The worked case: a percentile on a Counter Suppose someone copies a latency rule onto a request counter: ```javascript export const options = { thresholds: { http_reqs: ['p(95)<200'], }, }; ``` `p(95)<200` parses cleanly — the percentile argument is in range and the operator is legal — so stage one is happy. Stage two looks up `http_reqs`, finds a Counter, and asks whether a Counter supports the percentile method. It does not. k6 emits an error of the form *invalid threshold "p(95)<200" applied on metric http_reqs; reason: unsupported aggregation method p on metric of type counter. supported aggregation methods for this metric are: count, rate* and stops. Two details in that message are worth noticing. First, the method is reported as `p`, not as `p(95)`: the percentile argument is carried separately from the token, so the diagnostic names the family, not the instance. Second, the message enumerates the methods the type *does* accept, which is the fastest way to repair the rule without consulting the reference. ## Why the constraint exists Each type is backed by a different sink, and a sink can only answer the questions its shape allows. - A **Counter** stores one running total. It can tell you the total and the total divided by elapsed time; it retains no individual samples, so there is no distribution to take a percentile of. - A **Gauge** keeps only the last sample. Even though its internal representation happens to track the smallest and largest values it has seen, k6 publishes only `value` to the threshold layer, so `max<4000` on a Gauge is rejected. - A **Rate** keeps two counters. It can produce a proportion and nothing else. - A **Trend** retains every sample, which is why it is the only type that can answer `p(N)`. ## Practical consequences - Look up the metric's **type** before writing the expression; the metric's name tells you nothing. - A custom metric declared as `new Counter('errors')` gets the Counter vocabulary even if you conceptually think of it as a latency. - Because the check happens before the run, a mismatch costs you seconds in CI rather than a wasted load test — but only if the metric is registered at validation time. A metric created inside the default function rather than in the init context is not in the registry yet, and its threshold is rejected as a missing metric instead. - Re-declaring an existing metric name with a different type is itself an error, so you cannot widen a Counter into a Trend just to gain percentiles.
- Does a threshold on a k6 sub-metric such as `http_req_duration{status:200}` follow a different aggregation table?No. k6 resolves the key back to its parent metric and applies the parent's type, so a sub-metric of a Trend still accepts `avg`, `min`, `max`, `med` and `p(N)`, and a sub-metric of a Counter still accepts only `count` and `rate`. Narrowing by tag changes which samples are counted, never which aggregations are legal.
- What distinguishes a rejected aggregation from a breached threshold in k6?A rejected aggregation is a configuration fault found before any virtual user starts, so no test runs and no metrics are produced. A breached threshold is a test outcome: the run completes, the summary marks the rule with a cross, and k6 signals failure through its exit status. The first is fixed by editing the expression, the second by fixing the system under test.
- Why can a Gauge threshold not use `max`, given that k6's gauge sink tracks a maximum?The sink's minimum and maximum fields exist for internal bookkeeping and are never published to the threshold layer, which receives only `value` for a Gauge. Because the legal-method list is derived from the metric type rather than from what the sink happens to compute, `max<4000` on a Gauge is rejected as an unsupported aggregation.
saying these in an interview costs you the question
- Assuming any metric can take p(95) because percentiles are universal
- Expecting k6 to ignore an unsupported aggregation and keep running
- Thinking the metric name, not its type, decides the vocabulary
- Treating a rejected expression as a failed threshold result
- Believing a Gauge exposes min and max to thresholds