skip to content

In k6, which exit code wins when a run both breaches a threshold and marks itself failed?

level: seniorimportance: should knowfreq 34%

answer

  1. first outcome, not worst outcome
  2. the breach code is only a fallback
  3. codes never overwrite each other
  4. WithExitCodeIfNone and the err == nil guard

basics

~10 s

k6 exits 110, not 99. The threshold code is attached only if the run does not already carry an error, so exec.test.fail(), exec.test.abort() and script exceptions all take precedence over a breach.

solid answer

~40 s

The exit code reports **what k6 hit first**, not the worst outcome. `exec.test.fail()` makes the command return a marked-as-failed error as soon as the scheduler returns; thresholds are finalised afterwards, in a deferred step that assigns its own error only `if err == nil`. So the run exits **110** and logs `Crossed thresholds, but test already exited with another error`. The same precedence gives **108** for a self-abort and **107** for a script exception, while a malformed threshold expression exits **104** before any load runs. The breach is still computed and still shown in the summary — only the integer changes, which is why a non-99 code is no evidence that thresholds passed.

code

javascript · 16 lines
javascript
import http from 'k6/http';
import exec from 'k6/execution';

export const options = {
  vus: 1,
  iterations: 5,
  thresholds: { http_req_duration: ['p(95) < 1'] }, // certain to breach
};

export default function () {
  http.get('https://quickpizza.grafana.com');
  exec.test.fail('marked by the script');
}

// Both conditions hold. The process exits 110, and the summary
// still shows http_req_duration as breached.

go deeper

for a junior

Learn the ordering as a fact first: a script that aborts or marks itself failed reports its own code, and the threshold code appears only when nothing else did.

for a middle

Explain the mechanism: exit codes are attached only when the error carries none, and the threshold finaliser assigns its error only if the command has none yet.

for a senior

Draw the operational conclusion — a non-99 code is no evidence the rules passed, so a breach has to be read from the summary or log rather than inferred from the integer.

for a principal

The tradeoff to weigh is how much of the verdict should be expressed in measured rules versus in script control flow, since a hand-coded failure hides the breach behind a different number.

## 99 is the fallback, not the winner It is tempting to read k6's exit code as "the worst thing that happened". It is not. It is **the first thing that ended the run**, and a breached threshold is the last thing k6 gets around to deciding. When a run both breaches a threshold and hits something else, the something else wins. The concrete case: a script whose thresholds are certain to breach *and* which calls `exec.test.fail()`. The process exits **110**, not 99. ## Why, mechanically Two pieces of `internal/cmd/run.go` produce this. 1. **Thresholds are finalised in a deferred function.** After the run, k6 defers a `handleFinalThresholdCalculation` step that evaluates every threshold one last time and builds an error tagged `exitcodes.ThresholdsHaveFailed`. But it assigns that error to the command's return value only under a guard: `if err == nil { err = tErr }`. If an error is already there, k6 logs `Crossed thresholds, but test already exited with another error` at debug level and leaves the existing one alone. 2. **The code is attached with `WithExitCodeIfNone`.** Every exit code in k6 is attached to an error through a helper whose name says the rule out loud: it sets a code **only if the error does not already carry one**. Codes never overwrite each other. `exec.test.fail()` sets the run status; the command checks that status as soon as the scheduler returns and returns a `MarkedAsFailed` error. That happens before the deferred threshold step runs, so by the time thresholds finalise, the return value is already occupied. ## The ordering that results | what the run did | code the job sees | why | |---|---|---| | breached a threshold only | 99 | nothing else set a code | | breached and called `exec.test.fail()` | 110 | the marked-as-failed error was returned first | | breached and called `exec.test.abort()` | 108 | the interrupt came back from the scheduler as the run's error | | breached and threw an uncaught exception | 107 | the script exception is the run's error | | had a rejected threshold expression | 104 | configuration is rejected before any load runs | The last row is a different kind of "first": k6 parses and validates thresholds before the run starts, so a run with a malformed rule never generates a sample and never reaches the evaluation stage at all. ## What is *not* lost The breach itself is not discarded when another code wins. k6 still finalises thresholds, so: - the breached metrics are still marked in the end-of-test summary; - the breach is still visible in the log; - configured outputs still receive everything that was flushed. Only the integer changes. This is worth saying out loud because the natural reaction to "the job exited 110" is to assume the thresholds passed, and they may not have. Nor is this precedence an accident being worked around. It is exactly what the helper's name states: `WithExitCodeIfNone` sets a code only when the error does not already carry one, so the first mechanism to produce an error owns the number and every later mechanism defers to it by design. ## How to read a k6 exit code correctly The number answers one question: **what did k6 hit first?** It does not answer "was this the only problem" and it does not rank severity. Three consequences follow for anyone reading the code programmatically: 1. **Do not treat 99 as "the failure case" and everything else as an infrastructure problem.** A run that breached its rules can legitimately report 107, 108 or 110 instead. 2. **Do not infer from a non-99 code that thresholds passed.** Read the summary or the log for that. 3. **Do distinguish the codes that mean k6 never produced a verdict** — 104 for rejected configuration, 107 for a script exception — **from the codes that mean the run reached an end** — 99 and 110. That is a real difference in what the artefacts are worth. ## In k6 v2 The precedence behaviour follows from `WithExitCodeIfNone` and the `if err == nil` guard, both current in k6 v2.x. Nothing in the exit-code table changed in the v2 major line; what changed around it was the removal of the cloud-control subcommands and the REST API server no longer starting by default, so the `103` and `106` codes are now reachable only when `--address` is passed.

  • If a k6 run exits 108, can the team conclude its thresholds passed?
    No. 108 only says the script aborted the run itself, and abort takes precedence over the threshold code. k6 still finalises thresholds and still marks breached metrics in the summary, so the breach has to be read from the summary or the log rather than inferred from the integer.
  • Why does a malformed k6 threshold expression exit 104 rather than 99?
    Because it never becomes a verdict. k6 parses and validates every threshold expression before the run starts and wraps a failure in `exitcodes.InvalidConfig`, so no load is generated and no metric is ever evaluated. 99 requires a rule that ran and failed; 104 means k6 refused the rule.
  • Does the deferred threshold finalisation still run when another error already set the exit code?
    Yes. It evaluates every threshold a final time so the summary and outputs are correct, then finds the return value already occupied, logs `Crossed thresholds, but test already exited with another error` at debug level, and discards only its own error. The evaluation result is kept; the code is not.

saying these in an interview costs you the question

  • Assumes k6 returns the most severe code, not the first
  • Reads a non-99 exit code as proof thresholds passed
  • Thinks 99 overrides a script's own abort or fail
  • Believes a rejected threshold expression exits 99
  • Expects k6 to combine several outcomes into one code