skip to content

How does a failed k6 check() surface in the checks metric and the end-of-test summary?

level: middleimportance: should knowfreq 52%

answer

  1. one Rate metric, three summary rows
  2. one sample per check key
  3. zero for fail, one for pass
  4. the check tag carries the name
  5. total, succeeded and failed are derived

basics

~20 s

Each check key emits one sample on k6's built-in checks Rate metric, 1 for a pass and 0 for a fail, tagged with the condition's name. The summary derives checks_total, checks_succeeded and checks_failed from that single metric.

solid answer

~40 s

`checks` is a built-in **Rate** metric in k6 v2, and every key inside a `check()` object emits exactly one sample into it per evaluation: `1` when the condition is truthy, `0` when it is not. Each sample carries the system tag `check` set to that key's text — `check` is enabled by default — so a failure stays attributable to the individual condition. The summary never prints `checks` directly; it derives three display rows from that one Rate sink: `checks_total` (the sample count), `checks_succeeded` (the pass rate) and `checks_failed` (the fail rate), each rendered as a percentage plus an `N out of M` count. Those three names exist only in the summary and are not registered metrics.

code

javascript · 12 lines
javascript
import http from 'k6/http';
import { check } from 'k6';

export default function () {
  const order = http.post('https://example.com/checkout', '{"cart":"c-42"}');

  // Two keys -> two samples on the built-in `checks` metric, every iteration.
  check(order, {
    'checkout returned 201': (r) => r.status === 201,
    'order id present': (r) => r.json('orderId') !== null,
  });
}

go deeper

for a junior

Know where check results end up: one built-in metric called checks, and three rows in the end-of-test summary named checks_total, checks_succeeded and checks_failed.

for a middle

Explain the mechanics: one sample per key, valued 1 or 0, tagged with the condition's name, landing in a Rate sink that stores only a total and a count of trues. The three summary rows are arithmetic on those two numbers.

for a senior

Demonstrate that you can act on the data. Say how you would break the rate down by the check tag in a streamed output to find which condition degraded, rather than reading a single suite-wide percentage.

for a principal

Consider the cost of check granularity across a team's suites: more named conditions give sharper attribution through the check tag, but they also inflate cardinality in whatever backend the samples are streamed to.

## One metric, three summary rows k6 v2 ships a fixed set of built-in metrics, and exactly one of them carries check results: **`checks`**, a **Rate** metric. A Rate metric is the simplest sink k6 has — it keeps two counters, a total and a number of "trues", and its value is the quotient of the two. Everything you ever learn about how your checks behaved comes out of those two integers. The end-of-test summary does not print `checks` under that name. It expands the single sink into three display rows, which is what makes readers think three metrics exist: | Row in the summary | Shown as | Computed from the `checks` sink | | --- | --- | --- | | `checks_total` | a count, plus a per-second rate | the sink's total number of samples | | `checks_succeeded` | a percentage and `N out of M` | trues divided by total | | `checks_failed` | a percentage and `N out of M` | (total minus trues) divided by total | Those three names are **display values only**. They are not registered metrics, they are not emitted to output backends, and a threshold that names one is rejected before the test starts. The metric you can threshold, stream and tag is `checks`. ## What one failed condition writes The unit of measurement is the **key**, not the call. When k6 evaluates a `check()` object it loops over the object's keys and, for each one: 1. Resolves the value — invoking it with the first argument when it is a function. 2. Coerces the result to a boolean. 3. Emits a single sample on `checks`, valued `1` when the boolean is true and `0` when it is false. 4. Attaches the system tag `check`, set to that key's text, provided `check` is in the run's enabled system tags — and it is, by default. So a call with two keys running over 20 iterations produces 40 samples, not 20. A `checks_total` of 40 with `checks_succeeded` at 50% means twenty conditions were satisfied and twenty were not; it says nothing yet about *which* ones. ## Keeping the failure attributable The `check` tag is what stops a run's checks from collapsing into one anonymous percentage: - In the summary, k6 lists each distinct check name with its own pass and fail counts, so a checkout condition that fails every time is visible next to the ones that never do. - In a streamed output, every sample carries `check: <name>` alongside the usual scenario and group tags, so a dashboard can break the rate down by condition. - Because `check` is an ordinary, indexable system tag, a tagged threshold selector can single one condition out rather than scoring the whole suite. Two tags are the exception to that last point: `vu` and `iter` are non-indexable in k6, so selectors built on them warn and do not work. `check` is not one of them. ## The checkout example end to end An iteration posts a cart, then checks two things about the response — that the status was `201` and that an order id came back. Suppose the endpoint starts answering `500` under load. Each iteration now emits one sample valued `1` (nothing, in this case) and two valued `0`, both tagged with their respective check names. After 20 iterations the summary reads something like: ``` checks_total.......................: 40 3.94/s checks_succeeded...................: 0.00% 0 out of 40 checks_failed......................: 100.00% 40 out of 40 ``` Every one of those numbers came out of a Rate sink holding `total = 40, trues = 0`. And the run still exits `0`, because nothing here is a verdict — that is the job of a threshold on `checks`. ## Traps worth naming in an interview - **`checks_total` counts conditions, not calls.** Padding a single `check()` object with extra keys inflates it. - **A passing condition is recorded too.** k6 needs both the numerator and the denominator; it does not only write failures. - **`checks_failed` is not thresholdable.** Point the rule at `checks`; naming a display row fails configuration validation with `no metric name "checks_failed" found`. - **`checks` is a Rate, not a Counter.** That is why the summary can show a percentage at all, and why the only sensible thing to assert about it is a rate. - **Disabling the `check` system tag costs you attribution.** The totals survive, but every sample becomes indistinguishable and per-condition breakdown disappears.

  • Why can a k6 threshold not be written against checks_failed?
    `checks_failed` is a display value k6 computes from the `checks` Rate sink for the summary, not an entry in the metric registry, and it is not emitted to output backends either. Threshold validation looks the name up and aborts with `no metric name "checks_failed" found`. Target `checks` instead.
  • Does checks_total count check() calls or individual conditions in k6?
    Conditions. k6 emits one sample per key in every `check()` object it evaluates, so one call with three keys running for 100 iterations yields a `checks_total` of 300 rather than 100. The call itself contributes nothing beyond its keys.
  • How do you tell which k6 check failed when a run has dozens of them?
    Read the `check` system tag. Every sample on the `checks` metric carries `check: <name>` by default, so a streamed output or a tagged view separates the conditions, and the end-of-test summary lists each check name with its own pass and fail counts.

saying these in an interview costs you the question

  • Thinks checks_failed is a real metric you can threshold
  • Says checks_total counts check() calls rather than conditions
  • Believes a failing condition emits no sample at all
  • Assumes the check name is lost once the run aggregates
  • Treats checks as a Counter rather than a Rate metric