skip to content

Even with `set -euo pipefail` at the top, experienced bash authors still write `command || die "message"`. What does an explicit die helper add, and what does a good one look like?

level: middleimportance: should knowfreq 45%

answer

  1. errexit stops, it never explains
  2. message, stream, exit code
  3. stderr keeps stdout clean
  4. preflight checks before the work
  5. zero status is only a tool's opinion

basics

~20 s

An explicit die helper turns a silent abort into a diagnostic: it prints a named, actionable message to stderr and exits with a chosen status. It also covers cases set -e cannot see, such as a missing tool, a missing variable, or a command that reports failure while exiting zero.

solid answer

~50 s

`set -e` aborts, but it aborts silently and only on a non-zero exit status. A `die` helper adds three things. First, a message: `printf '%s: %s\n' "${0##*/}" "$*" >&2` tells the operator which script failed and why, on stderr so it does not pollute the script's real output. Second, a chosen exit code, so the caller — a CI step, a systemd unit, a parent script — can distinguish a usage error from a runtime failure. Third, coverage of failures errexit never sees: a required tool that is absent (`command -v jq >/dev/null || die "jq is required"`), a required variable that is empty, or a tool such as `curl` without `--fail` that exits 0 on an HTTP error. The `||` form also works in the contexts where errexit is suspended, which is why it is a habit rather than a fallback.

code

bash · 15 lines
bash
#!/usr/bin/env bash
set -euo pipefail

die() {
  printf '%s: %s\n' "${0##*/}" "$*" >&2
  exit 1
}

command -v jq >/dev/null 2>&1 || die "jq is required but not installed"
[[ -n "${API_URL:-}" ]] || die "API_URL is not set"

curl -fsS "$API_URL/manifest" -o manifest.json || die "could not fetch manifest from $API_URL"
[[ -s manifest.json ]] || die "manifest downloaded but is empty"

jq -r '.version' manifest.json

go deeper

for a junior

Know that error messages belong on stderr and that a failing script must exit non-zero. Being able to write a three-line die and call it with || already puts a script ahead of most throwaway ones.

for a middle

Justify each part of the helper: printf over echo, ${0##*/} for the prefix, >&2 so piped stdout stays clean, and a deliberate exit code. Explain why || die behaves the same with or without set -e.

for a senior

Focus on the failures strict mode cannot see — missing dependencies, empty-but-set configuration, tools that exit zero on a logical failure — and on asserting the outcome you wanted rather than trusting the tool's status.

for a principal

Treat exit codes as an interface. Decide what codes the scripts in your estate publish, how the callers that consume them branch, and where a shared shell library of helpers is worth the coupling it introduces.

## What errexit leaves on the table Strict mode buys you a stop. It does not buy you an explanation, a code, or coverage of anything that fails without exiting non-zero. Consider a deploy script that dies at line 60 under `set -e`. The log shows whatever that command printed — possibly nothing — and the shell contributes silence. Compare with `… || die "could not reach the artifact store at $url"`, which names the operation and the input. There are three distinct gaps: 1. **No diagnostic.** Errexit prints nothing of its own. 2. **No control over the exit code.** The script exits with whatever the failing command returned, so `2` might mean "usage error" from your own script or "file not found" from `grep`. 3. **No coverage of non-status failures.** Errexit sees exit codes only. ## A serviceable die ```bash die() { printf '%s: %s\n' "${0##*/}" "$*" >&2 exit 1 } ``` Every element earns its place: - **`printf` rather than `echo`.** `echo` with flags such as `-e` is not portable across shells and builds, and `printf` with an explicit format string never mistakes a message beginning with `-` for an option. - **`${0##*/}`** strips the directory, so the message is prefixed with just the script's name — vital when several scripts write to the same log. - **`"$*"`** joins the arguments into one message, so `die "cannot read" "$file"` still produces one line. - **`>&2`** sends it to standard error. This is not cosmetic: if the script's stdout is being consumed by a pipeline or captured into a variable, an error message on stdout silently corrupts the data. - **`exit 1`** guarantees a non-zero status. A variant taking the code as the first argument, `die 2 "usage: ..."`, lets callers distinguish categories. ## Where you call it ```bash #!/usr/bin/env bash set -euo pipefail die() { printf '%s: %s\n' "${0##*/}" "$*" >&2; exit 1; } command -v jq >/dev/null 2>&1 || die "jq is required but not installed" [[ -n "${API_URL:-}" ]] || die "API_URL is not set" curl -fsS "$API_URL/manifest" -o manifest.json || die "could not fetch the manifest from $API_URL" [[ -s manifest.json ]] || die "manifest is empty" ``` Notice what each check catches that strict mode does not: - **Preflight checks.** A missing dependency is not a failing command; it is a command that never runs until line 80, where it produces `jq: command not found` and status 127. Checking it up front converts a confusing late failure into a clear early one. - **Required configuration.** `set -u` fires only when the variable is *unset*; a variable exported as the empty string passes it. `[[ -n "${API_URL:-}" ]]` catches both. (The `:-` is needed so the test itself does not trip nounset.) - **Success that is not success.** `curl` without `--fail` exits 0 on an HTTP 404 and writes the error page to the output file. The `-f` flag fixes that specific case, but the general lesson is that exit status is a tool's *opinion*, and for some tools it is a poor one. `[[ -s manifest.json ]]` asserts the outcome you actually wanted. ## Why `||` and not `if` `cmd || die "..."` reads as a single thought and stays on one line, which matters when a script has thirty of them. It also works everywhere: because errexit is suspended for the left-hand side of a `||` list anyway, this form behaves identically whether or not `set -e` is active — one less thing that changes meaning depending on the preamble. When you need both the output and the check, use the `if` form instead, so the value is not lost: ```bash if ! out=$(some_command); then die "some_command failed" fi ``` ## Exit codes as a contract Once a script has a `die`, giving it a code argument is nearly free, and it lets the script participate in a contract: 0 success, 1 general runtime failure, 2 usage error, and a small set of documented codes for conditions the caller genuinely branches on. Whatever the scheme, document it in the script's header comment — an undocumented code is only marginally better than an undifferentiated 1. ## The honest limit A `die` habit costs a line at every call site, and its value is proportional to how far the failure is from the person reading the log. For a five-line script you run by hand, `set -e` alone is fine. For something that runs at 3am under cron and whose only witness is a log file, every failure should name itself.

  • Why does a die helper use `printf` rather than `echo`?
    `echo`'s handling of flags and backslash escapes differs between shells and builds, so `echo -e` or a message starting with `-n` behaves inconsistently. `printf '%s\n' "$msg"` has one defined behaviour everywhere and never interprets the message as options. It also composes cleanly when you want a prefix, as in `printf '%s: %s\n' "$name" "$msg"`.
  • How would you extend `die` so callers can choose the exit code?
    Take the code as the first parameter and shift: `die() { local code=$1; shift; printf '%s: %s\n' "${0##*/}" "$*" >&2; exit "$code"; }`. Callers then write `die 2 "usage: ..."` for a usage error and `die 1 "..."` for a runtime failure. Document the scheme in the script header, because an undocumented code tells the caller nothing.
  • Your die helper is defined inside a function that runs in a command substitution. Does its `exit` stop the whole script?
    No. A command substitution runs in a subshell, so `exit` there terminates only that subshell; the parent sees a non-zero status from the substitution and continues, or aborts only if errexit applies at that point. This is a classic reason a script "ignores" its own die, and a reason to check the substitution's status in the caller.

saying these in an interview costs you the question

  • Writes error messages to stdout instead of stderr
  • Assumes exit status zero always means the command succeeded
  • Thinks set -e removes the need for any explicit checks
  • Uses echo -e in a helper and calls it portable
  • Exits with 0 after printing an error message

context