skip to content

How does git bisect run automate the search, and what do its exit codes mean?

level: middleimportance: must knowfreq 45%

answer

  1. the script answers, not you
  2. zero means good
  3. one special code means untestable
  4. 125 is skip
  5. 128 and up aborts

basics

~20 s

git bisect run <cmd> runs your test at each step and answers for you: exit 0 means good, 1 to 127 except 125 means bad, 125 means untestable so the commit is skipped, and 128 or above aborts the bisect.

solid answer

~40 s

After marking the endpoints you hand Git a command: `git bisect run ./check.sh`. Git checks out each candidate, runs the command, and converts its exit status into an answer — `0` means good (or old), any status from `1` to `127` except `125` means bad (or new), `125` means the commit cannot be tested so it is skipped, and anything else, `128` and above, aborts the session. The whole search then runs unattended and ends with `<sha> is the first bad commit`. In practice the script builds first and exits `125` if the build fails, then runs the narrowest reproducer that distinguishes the two states. Keep the script outside the working tree, or copy it to a temporary path, because checking out old commits can change or delete a script stored in the repository.

code

bash · 5 lines
bash
#!/bin/sh
# /tmp/check.sh - kept outside the repo so checkouts cannot change it
make -s >/dev/null 2>&1 || exit 125   # unbuildable: skip this commit
./build/app --parse fixtures/in.txt | grep -q EXPECTED || exit 1  # bad
exit 0                                # good

go deeper

for a junior

Know that git bisect run takes a command and answers each step for you, and that the command must exit zero when the code is good and non-zero when it is broken.

for a middle

Be precise on the contract: 0 good, 1 to 127 except 125 bad, 125 untestable, 128 and above aborts. Explain why 126 and 127 make setup errors look like real failures.

for a senior

Demonstrate script design judgment: build first and skip on failure, test one narrow deterministic symptom, keep the tree clean and the script outside the repo, and validate the script on both endpoints first.

for a principal

Argue for the investment that makes automated bisect routine — per-commit buildability, a fast deterministic reproducer, and a culture that records regressions as runnable checks rather than prose.

## From manual to automated A manual bisect makes you build, test and type an answer at every step. `git bisect run <cmd> [<args>]` replaces the typing: you still mark endpoints with `git bisect bad` and `git bisect good` (or `git bisect start <bad> <good>`), then Git checks out each candidate, executes the command, and derives the answer from its exit status. Ten steps that took a coffee break become a single unattended run. ## The exit-code contract This is the part interviewers probe, because it is a real protocol: - `0` — the commit is good (or old with custom terms). - `1` through `127`, except `125` — the commit is bad (or new). A test runner that exits non-zero on failure fits this naturally. - `125` — the commit cannot be tested; Git treats it exactly like `git bisect skip` and picks a nearby commit instead. This is the code for a broken build, not for a failing test. - `128` and above — abort the bisect entirely. Use this when the environment itself is wrong and continuing would be meaningless. Note the trap in the middle band: `126` and `127` are the shell's codes for a command that is not executable and a command that is not found. Because they fall in the bad range, a mistyped path or a script without the executable bit silently marks every commit bad and produces a confident, wrong answer. Sanity-check by running the script by hand on a known-good and a known-bad commit before starting. ## Writing the script A good bisect script is short, deterministic and fast, and it separates cannot judge from is broken: 1. Build. If the build fails, exit `125` — an unbuildable commit says nothing about the bug. 2. Reproduce one specific symptom, as narrowly as possible: a single unit test, a small script, a grep of output. A whole test suite is slow and can fail for unrelated reasons, which corrupts answers. 3. Exit `0` when the symptom is absent and `1` when it is present. It must also leave the tree clean. Bisect checks commits out, so build artefacts left as modifications to tracked files can block the next checkout; keep outputs untracked or clean them up. Similarly, avoid tests that depend on the network or on shared mutable state, since a flaky failure marks an innocent commit bad. ## Where the script lives Checkouts rewrite the working tree, so a script stored in the repository may change, vanish or become an older version as bisect walks. Keep it outside the repo and invoke it by absolute path, for example `git bisect run /tmp/check.sh`. The same applies to fixtures the script needs. ## Related mechanics `git bisect run` obeys custom terms, so with `--term-old` and `--term-new` set, exit `0` means old and non-zero means new. You can combine it with a path limit from `git bisect start -- <path>`. When the run finishes the session is still open: read the reported commit, then `git bisect reset`. If the run went wrong, `git bisect log` shows every answer it recorded, and `git bisect replay <file>` can re-drive a corrected session. ## Why it matters beyond the command Automated bisect is the point where commit hygiene turns into diagnosis speed. It only works if commits build individually and one command can decide the symptom — which is also why teams that squash a week of work into a single commit get little out of it.

  • Why exit 125 instead of just failing when the commit does not build?
    Because a failed build is not evidence about the bug. Exiting 1 would mark the commit bad and could send the search into the wrong half, converging on an innocent commit. Exit 125 tells Git the commit is untestable, so it picks a neighbour and keeps the remaining range honest.
  • Your bisect run reports the very first commit as bad. What do you check?
    Verify the script by hand on both endpoints. Common causes are exit 127 because the script path is wrong or not executable, a script that always returns non-zero due to a missing fixture, or endpoints marked the wrong way round. Both 126 and 127 fall in the bad range, so setup errors look like genuine failures.
  • How do you keep a slow test suite from making bisect run impractical?
    Reduce the test to the narrowest reproducer that still distinguishes the two states — one test case or a small script rather than the full suite — and use an incremental build. Limiting the search with git bisect start -- <path> shrinks the candidate set further when you know which subsystem is involved.

saying these in an interview costs you the question

  • Says any non-zero exit code means the commit is bad
  • Uses exit 1 for a broken build instead of 125
  • Keeps the bisect script inside the repository being checked out
  • Runs a whole flaky test suite as the bisect test
  • Thinks bisect run needs no good and bad endpoints

context