skip to content

Why can a k6 run gate a CI pipeline with no extra tooling, and where does that native gate stop?

level: seniorimportance: should knowfreq 55%

answer

  1. the tool renders the verdict itself
  2. non-zero exit is the whole integration
  3. breached threshold exits 99
  4. no threshold declared, no verdict
  5. a failed check never fails the run

basics

~20 s

k6 declares its pass rule in the script's exported options and exits 99 when a threshold is breached, so the CI step is just the run. But a failed check() never changes that code, and a script with no thresholds exits zero.

solid answer

~50 s

In k6 v2 the pass rule ships in the script itself, under `thresholds` in the exported `options` object, and a breach makes the k6 process exit **99**. Every CI runner already fails a step on a non-zero exit, so the pipeline stage is literally `k6 run script.js` — no report parser, no wrapper that greps output, no separate assertion job. Three limits matter. The gate is threshold-shaped, so it judges metrics and nothing else. A failed `check()` returns `false` but never changes the exit code, so a run of failing checks still exits zero unless a threshold covers the checks metric. And a script declaring no thresholds always exits zero — k6 will happily run a test with no verdict. Non-zero is also not one thing: an invalid configuration or an uncaught exception exit with their own codes.

code

bash · 2 lines
bash
k6 run script.js
echo "k6 exited $?"   # 99 when a declared threshold was breached, 0 when none were

go deeper

for a junior

Learn the one mechanic: a breached threshold makes k6 exit 99, and CI fails a step on a non-zero exit. That is why the pipeline stage can be nothing more than the run itself.

for a middle

Be able to say what the gate does not cover — checks do not change the exit code, no threshold means no verdict, and other failures also exit non-zero with their own codes.

for a senior

Show you have been burned: describe how you would prove a green stage is meaningful, and how you would stop a broken script from being reported as a breached threshold.

for a principal

Argue the selection point. A built-in verdict removes a class of glue code permanently, but it shifts the whole risk onto whether teams declare thresholds at all — decide how you enforce that.

## The verdict lives in the script k6 carries its own pass rule. The exported `options` object has a `thresholds` key, and when the run ends with one of those thresholds breached, the k6 process **exits 99** — the code k6 reserves for "thresholds have failed". Nothing else in the toolchain has to be involved. That single fact is what makes k6 cheap to put in a pipeline. Every CI system on the market already fails a step when the command it ran exits non-zero, so the integration is: ```bash k6 run script.js ``` There is no results file to parse, no wrapper script that greps stdout for a word, no second job that reads a report and decides. The tool that generated the load also rendered the verdict, and it rendered it in the one currency the pipeline already speaks. ## Why that matters when choosing a generator - **The gate is not something you build.** Teams often spend their first week with a load tool writing the glue that turns its output into a build result. With k6 that step does not exist. - **The rule travels with the test.** Because the threshold is in the same module as the request code, you cannot change what "passing" means without it showing up in the same diff. - **Local and CI agree.** The same command on a laptop returns the same exit code, so a red build is reproducible without pipeline access. ## Where the native gate stops - **It is threshold-shaped.** Thresholds are evaluated over metrics, so anything that is not expressed as a metric cannot be gated this way, no matter how important it is. - **`check()` is not a gate.** In k6 a `check()` call returns a boolean and **never changes the exit code**; a condition that merely evaluates false does not abort the iteration. (A predicate that itself throws is the exception — the error propagates out of `check()`.) A run whose checks all failed still exits zero on its own. Only a threshold declared over the built-in checks metric turns that into a build failure. - **No thresholds means no verdict.** k6 does not supply a default threshold. A script that declares none exits zero every time, and a green pipeline stage then means only that the process ran to completion. - **Non-zero is not one thing.** An invalid configuration and an uncaught script exception each exit with their own distinct code. A pipeline that only asks "was the exit code zero?" will report a syntax error and a genuine breach identically. - **The run's own output is a terminal summary.** k6 prints an end-of-test summary; a dashboard, a file, or a time-series backend comes from an output sink or the summary hook, not from the run by default. ## Two failures that look the same and are not | what happened | exit code | what a naive CI step reports | |---|---|---| | a declared threshold was breached | 99 | build failed | | the script threw an uncaught exception | its own non-zero code | build failed | | every `check()` failed, no threshold declared | 0 | build passed | | no thresholds declared at all | 0 | build passed | The first two rows are why "non-zero" is not a diagnosis, and the last two are why a green stage is not evidence. ## Making the gate mean something 1. **Require a threshold.** Treat a k6 script with an empty or absent `thresholds` block as an incomplete test in review, because the tool will not complain about it. 2. **Distinguish the codes.** If the pipeline reports anything more than pass/fail, branch on the actual exit code so a broken script is not reported as a breached threshold. 3. **Cover checks deliberately.** If correctness assertions must fail the build, declare a threshold over the checks metric; otherwise be explicit that checks are diagnostic output only. ## The honest summary k6's native gate is a genuine differentiator when you are choosing a generator: the verdict is a first-class feature of the binary rather than something the team assembles from report files and shell. What it is not is a general-purpose assertion engine. It judges metrics against declared thresholds, and it declares nothing on your behalf. The work it saves you is the plumbing; the work it leaves you is deciding what to assert and remembering to assert it at all.

  • Your k6 stage has been green for a month. What would you check before trusting it?
    That the script actually declares thresholds. k6 supplies no default, so a test with none exits zero on every run regardless of what it measured. A green stage backed by an empty thresholds block is reporting only that the process finished.
  • How would you keep a broken k6 script from being reported as a performance regression?
    Branch on the exit code rather than on zero-versus-non-zero. A breached threshold exits 99; an invalid configuration and an uncaught script exception each have their own codes. Mapping them to different pipeline messages stops a syntax error from being reported as a breached threshold.

saying these in an interview costs you the question

  • Believing a failed check() makes k6 exit non-zero
  • Assuming k6 applies a default threshold when the script declares none
  • Treating every non-zero k6 exit as a performance failure
  • Thinking a report file must be parsed to fail a k6 build step
  • Saying a green k6 stage proves the run met a performance target