A bash script is run by a CI job that treats a non-zero exit as a build failure. If the script never calls `exit`, what determines the status it returns, and which conventions should you follow when choosing codes?
answer
- nothing explicit means the last command
- zero is success, everything else is not
- a trailing echo hides the failure
- one byte only, values wrap
- 127 and 128+N are already taken
basics
~20 sA bash script that ends without calling exit returns the exit status of the last command it ran, so a trailing echo or log line reports success even after an earlier failure. Zero means success, any non-zero value means failure, and the status is truncated to the range 0-255.
solid answer
~50 sWhen a script runs off the end of the file, its exit status is that of the last command executed — and `exit` with no argument does the same. That is a real bug source: a script whose final line is `echo "done"` or a cleanup command reports success no matter what happened earlier, which is why CI jobs sometimes go green on a broken build. Conventionally 0 means success and every non-zero value means failure; 2 is often a usage error, while the shell itself produces 126 for a file that is not executable, 127 for a command not found, and 128+N when a command was terminated by signal N. The status is a single byte, so `exit 300` actually reports 44 and `exit -1` reports 255. Choose a small set of codes, document them in the script header, and make sure the last thing the script does reflects the outcome.
code
bash · 4 lines#!/usr/bin/env bash
# Reports success even though grep found nothing.
grep -q ERROR app.log
echo "scan complete"go deeper
Say plainly that 0 means success and anything else means failure, and that a script with no exit returns the status of its last command. Recognise the trailing-echo bug when you see it.
Explain the conventional codes — 1 general failure, 2 usage, 126, 127, 128+N — and the single-byte truncation that turns exit 300 into 44. Show how to capture $? immediately and propagate a wrapped tool's status.
Treat the code set as an interface: pick a small documented range, avoid values the shell already claims, and make sure a script that aborts under strict mode still exits with a status the CI job or service manager can act on.
Own the convention across the estate. Define which codes callers may branch on, which ones signal a retryable condition, and how that contract stays stable as scripts are rewritten or replaced by programs in other languages.
## What a script returns Every process ends with a small integer status that its parent can read. For a shell script that status is set by one of three things: 1. An explicit `exit N`. 2. An `exit` with no argument, which uses the status of the most recently completed command. 3. Running off the end of the file, which is identical to case 2. So the default is *the last command's status*, and that is what makes the following script lie: ```bash #!/usr/bin/env bash grep -q ERROR app.log # exits 1 when nothing matches echo "scan complete" # this is the last command; it succeeds # script exits 0 ``` A CI job running this sees success. The same trap catches scripts whose final line is a cleanup step, a log write, or a summary `printf`. Under `set -e` an *aborting* failure does propagate correctly — the shell exits with the failing command's status — but a failure in a position where errexit does not fire, followed by a successful last line, still produces 0. The mirror-image bug also exists: a script whose final line is a test, such as `[[ $count -gt 0 ]]`, returns 1 whenever the count is zero, failing a build for no reason. Ending with an explicit `exit 0` when the outcome really was success removes both surprises. ## The conventions The shell's own contract is minimal and worth stating precisely: - **0 means success. Everything else means failure.** There is no "partial success" code and no "warning" code that a caller will interpret for you. - **1** is the general-purpose failure. Most tools use it for "the thing you asked for did not work". - **2** is widely used for a usage error — bad arguments, missing required option. Many GNU tools follow this, and `getopts`-driven scripts commonly adopt it. - **126** — the command was found but could not be executed (typically not executable). - **127** — command not found. Seeing 127 from a script almost always means a missing dependency or a `PATH` that is not what you assumed. - **128+N** — the command was terminated by signal N, so 130 for an interrupt (signal 2) and 143 for a termination request (signal 15). This is how bash reports a signalled child; the encoding is the shell's reporting convention. Because these are conventions rather than enforcement, nothing stops a script from returning 127 for its own reasons — but doing so guarantees someone will misdiagnose it as a missing binary. ## The range is one byte The status a parent can read is masked to 8 bits. Two consequences bite in practice: ```bash exit 300 # reported as 44 (300 - 256) exit -1 # reported as 255 exit 256 # reported as 0 — a failure that reads as success ``` Any scheme that maps, say, an HTTP status code straight onto an exit code is broken by this. Keep codes in 1–125 for your own conditions, leaving the upper range to the shell's conventions. ## Designing the contract When the caller genuinely branches on the reason for failure — a CI step that retries on a transient error but not on a validation error, or a service manager configured to restart only on certain outcomes — a small documented code set is worth having: ```bash # Exit codes: # 0 success # 1 runtime failure # 2 usage error # 3 configuration invalid # 4 upstream unavailable (safe to retry) ``` Three rules keep it honest. Keep the set small, because every code is an interface the caller must handle. Document it in a header comment, because an undocumented 3 is no more informative than 1. And never reuse a code the shell already owns. ## Checking the status of what you just ran Interactively and in scripts, `$?` holds the status of the last completed command, and it is overwritten by the *next* command — including a test or an `echo`. If you need it more than once, save it immediately: ```bash some_command status=$? if (( status != 0 )); then printf 'some_command failed with %d\n' "$status" >&2 exit "$status" fi ``` Propagating the original status rather than a flat `exit 1` preserves whatever information the underlying tool encoded, which is usually the most useful thing a wrapper script can do.
- Your script wraps another tool and you want the caller to see that tool's exit code. How do you do it?Capture the status immediately after the tool runs, because `$?` is overwritten by the next command: `some_tool "$@"; status=$?`. Then, after any logging or cleanup, finish with `exit "$status"`. Propagating the original value keeps whatever meaning the tool encoded, whereas a flat `exit 1` throws that information away.
- Why does a script that ends with `[[ -s report.txt ]]` sometimes fail a CI job that had no real error?The test is the last command, so its status becomes the script's status. When the file is empty the test returns 1, and the job reads that as a failure even though nothing went wrong. Any script whose final statement is a conditional should end with an explicit `exit 0` on the success path.
- What exit status does a bash script report when the command it was running is killed by a termination signal?Bash reports 128 plus the signal number, so a process terminated by signal 15 shows 143 and one interrupted by signal 2 shows 130. That encoding is how the shell surfaces "this did not exit on its own" through a single-byte status, and it is why your own codes should stay well below 128.
saying these in an interview costs you the question
- Thinks a script always exits 0 unless exit is called
- Assumes the first failing command sets the script's status
- Returns a value above 255 as a distinct code
- Uses exit 127 for an application-specific error
- Reads $? after running another command in between