skip to content

How does k6 arrive at the p(95) figure it prints for a Trend metric?

level: middleimportance: should knowfreq 47%

answer

  1. the samples are all still there
  2. sorted once, lazily
  3. rank scaled by n minus one
  4. straddling values, split proportionally
  5. med is the same call with 0.5

basics

~20 s

k6's Trend sink keeps every value the run recorded. For p(95) it sorts them, computes the rank 0.95 * (n - 1), and returns the value at that index, interpolating linearly between the two values that straddle it.

solid answer

~40 s

A k6 `Trend` sink appends every observed value to a slice and keeps a running count, sum, minimum and maximum. When the summary needs `p(95)`, the sink sorts the slice once, converts the token to a fraction, and computes the rank `0.95 * (n - 1)`. If that rank is a whole number it returns the value at that index; otherwise it interpolates linearly between the values at `floor` and `ceil` of the rank. `med` is the same routine called with `0.5`, `p(0)` returns the minimum and `p(100)` the maximum, while `avg`, `min`, `max` and `count` come straight from the running fields and need no sort. There is no estimator setting and no bucket layout to configure on a k6 Trend.

code

javascript · 12 lines
javascript
import { Trend } from 'k6/metrics';

const latency = new Trend('checkout_latency', true);

export const options = {
  summaryTrendStats: ['avg', 'min', 'med', 'max', 'p(90)', 'p(95)', 'p(99)', 'count'],
};

export default function () {
  // Each add() appends one value to the sink that p(99) is later read from.
  latency.add(123);
}

go deeper

for a junior

Know that a k6 Trend holds the individual values and that percentile columns are read from them, so the summary figures are not approximations of a compressed distribution.

for a middle

Walk through the rank formula and the interpolation between the two straddling values, and note that avg, min, max and count come from running fields rather than the sorted slice.

for a senior

Connect retention to memory: one sink per metric plus one per threshold sub-metric, each growing with sample count, and know the 100,000 time-series warning that flags high-cardinality tags.

for a principal

Consider how long-running or high-cardinality k6 tests should be shaped so retained samples stay affordable on the generator, and what that implies for splitting a suite.

## Where the number comes from Every `Trend` metric in k6 is backed by a sink that stores the individual observations. On each sample the sink appends the value to a slice and updates three running fields — a count, a running sum, and the smallest and largest values seen. Nothing is bucketed, quantised or thrown away, so at the end of the run the sink is holding the complete set of values recorded for that metric. That is what makes the percentile calculation direct. When the summary needs `p(95)`, k6 does not consult a histogram or an estimator; it reads the figure off the values it kept. ## The algorithm, step by step 1. If the slice is not already sorted, sort it ascending, and mark it sorted so later percentile reads on the same sink skip the work. 2. Convert the requested percentile to a fractional rank. The token `p(95)` becomes `0.95`, and k6 computes `i = 0.95 * (n - 1)` where `n` is the number of recorded values. 3. If `i` is a whole number, the value at that index is the answer. 4. Otherwise take the two values that straddle it — the ones at `floor(i)` and `ceil(i)` — and interpolate linearly between them by the fractional part of `i`. Two shortcuts sit in front of that: an empty sink returns `0`, and a sink holding exactly one value returns that value for any percentile. Worked through with a tiny sink, the arithmetic is easy to follow. Suppose five values were recorded and sort to `[100, 120, 140, 200, 900]`, so `n = 5`. For `p(95)`, the rank is `0.95 * 4 = 3.8`. That is not a whole number, so k6 takes the values at index 3 and index 4 — `200` and `900` — and moves 0.8 of the way from the first to the second: `200 + (900 - 200) * 0.8 = 760`. Note that `760` is not a value any sample actually had. Interpolation means a k6 percentile can be a number that sits between two observations, and that is expected rather than a bug. The sorted flag matters for a report that asks for several percentiles. The first percentile read sorts the slice and records that it is sorted; `p(90)`, `p(95)` and `p(99)` in the same summary then reuse that ordering. The flag is cleared whenever a new value is added, so a sink that is still receiving samples re-sorts on its next read. ## What each token costs to compute | token | how k6 produces it | |---|---| | `avg` | running sum divided by the count; needs no sorted slice | | `min` / `max` | running fields maintained on every `Add` | | `count` | the running count of recorded values | | `med` | the percentile routine called with `0.5` | | `p(N)` | the percentile routine called with `N / 100` | So `med` and `p(50)` are literally the same call, `p(0)` lands on index `0` and returns the minimum, and `p(100)` lands on index `n - 1` and returns the maximum. ## What k6 does not give you here There is no estimator to select and no bucket layout to configure on a k6 `Trend`. The only knobs on this surface are which columns to print (`summaryTrendStats`) and which unit to print them in (`summaryTimeUnit`). If you have seen a percentile configuration option in another tool and go looking for its k6 equivalent, there is not one — the sink's contents *are* the configuration. ## The cost of keeping every value Retention is not free, and it is worth knowing where the memory goes in a long run: - k6 keeps **one sink per metric**, and a sink's slice grows by one `float64` for every sample the metric receives. - It also keeps **one sink per threshold sub-metric** — a threshold written against a tagged selector creates its own sink, and every matching sample is added to that sink as well as to the parent's. - Nothing is trimmed mid-run: the values are needed at the end, so they are held for the whole run. - The number of distinct time series a run produces is tracked separately, and k6 warns once it passes 100,000, pointing at high-cardinality tag values as the usual cause. The practical consequence is that a `Trend`'s memory footprint scales with the number of samples it takes, not with how many statistics you ask the summary to print. Adding or removing columns changes what is rendered, never what is retained. ## Why this matters when reading a k6 report Because the figures are read off retained values, `min`, `max` and any percentile in a k6 summary all come from the same sorted set of observations for that metric, and they are internally consistent with one another and with the `count` column. When you need to reason about a k6 number, the question to ask is which samples reached that sink — which scenario, which tag selector, which sub-metric — rather than how the statistic was approximated.

  • In k6, does asking for more percentile columns make the run slower or heavier?
    Barely. The values are retained regardless of which columns you list, and the slice is sorted once and reused for every percentile read on that sink. Each extra token is one more lookup into an already-sorted slice at report time, so the cost is in retention, not in the column list.
  • What does a k6 Trend sink return for p(95) when only one value was recorded?
    That single value, for every percentile. The sink short-circuits: a count of zero returns 0 and a count of one returns the only value, so `min`, `med`, `max` and every `p(N)` are identical. It is a useful tell that a metric received almost no samples.
  • Does a tagged k6 sub-metric get its own percentile, or share the parent's?
    Its own. A threshold written against a tagged selector creates a sub-metric with its own Trend sink, and only samples whose tags match are added to it. Its percentiles are therefore computed from a different, smaller set of values than the parent metric's, and the two figures will not agree.

It is the spreadsheet method: sort the column, walk 95% of the way down it, and if you land between two rows, take the proportional point between them.

saying these in an interview costs you the question

  • Saying k6 estimates percentiles from histogram buckets
  • Claiming k6 keeps only a rolling window of samples
  • Thinking p(95) picks the nearest sample without interpolating
  • Believing a Trend's memory depends on the column list
  • Assuming med and p(50) are computed differently