A bash deploy script that parses flags with getopts kept running with its default settings when a user passed a mistyped flag, and separately a CI step that runs `./deploy.sh -h` is marked as a failed build. What is wrong with the script's option-parsing contract, and how should it signal a usage error versus a help request?
answer
- a bad flag does not stop the loop
- strict mode cannot see this failure
- help is a success, not an error
- two streams, two exit codes
- one usage function, two call sites
basics
~20 sAn invalid option does not end a getopts loop — getopts still returns success — so a script with no error branch runs on defaults. Usage errors belong on stderr with a non-zero status such as 2; a -h help request belongs on stdout with exit 0.
solid answer
~50 sTwo contract bugs. The mistyped flag survived because an invalid option does not terminate a getopts loop: getopts sets the variable to `?` and still returns success, so without an explicit `\?)` branch that prints usage and exits, the loop just moves on and the script proceeds with its defaults — `set -e` cannot help, because nothing failed. The CI failure is the mirror image: `-h` is a *successful* request, so its output belongs on stdout with exit status 0, while a usage error belongs on stderr with a non-zero status, conventionally 2 to distinguish "you invoked me wrong" from "the work failed" (which is 1). Make the usage function write to stdout and have the error branches redirect it with `usage >&2` before exiting 2. Then a wrapper can pipe `-h` into a pager and CI can trust the exit code.
code
bash · 33 lines#!/usr/bin/env bash
# exit 0 ok | 1 operation failed | 2 usage error
set -euo pipefail
usage() {
cat <<EOF
Usage: ${0##*/} [-v] [-f FILE] TARGET
-f FILE write output to FILE
-v verbose output
-h show this help
EOF
}
verbose=0
outfile=""
while getopts ":hf:v" opt; do
case $opt in
h) usage; exit 0 ;;
f) outfile=$OPTARG ;;
v) verbose=1 ;;
:) printf '%s: option -%s requires a value\n' "${0##*/}" "$OPTARG" >&2
usage >&2; exit 2 ;;
\?) printf '%s: unknown option -%s\n' "${0##*/}" "$OPTARG" >&2
usage >&2; exit 2 ;;
esac
done
shift $((OPTIND - 1))
if [[ $# -ne 1 ]]; then
printf '%s: exactly one TARGET is required\n' "${0##*/}" >&2
usage >&2; exit 2
figo deeper
Remember that the case block needs a branch for unrecognised options, and that it should print usage and exit rather than fall through.
Explain the mechanism: getopts returns success on an invalid option, so the loop continues and strict mode sees no failure; the branch is the only guard.
Demonstrate the caller's view — help on stdout with 0 so wrappers and CI can use it, usage errors on stderr with a distinct non-zero status so a bad invocation is distinguishable from a failed run.
Fix the exit-status vocabulary across the estate and require it in review, so automation can branch on status without scraping messages and no tool can silently proceed after an unrecognised flag.
## Bug one: an invalid option is not the end of the loop The most common misreading of `getopts` is that a bad option stops it. It does not. On an unrecognised option, getopts assigns `?` to your variable (and, in silent mode, the offending letter to `OPTARG`) and **returns success**, so the `while` condition is still true and parsing continues with the next argument. Only exhausting the arguments, or reaching an operand or `--`, ends the loop. So a `case` with branches for the known letters and nothing else treats `--dryrun` or `-x` as noise. The script runs with its defaults. Strict mode does not save you: `set -e` fires on a command that returns non-zero and getopts returned zero; `set -u` fires on an unset variable and every variable was initialised. The only thing that turns a typo into a refusal is an explicit branch: ```bash \?) printf '%s: unknown option -%s\n' "${0##*/}" "$OPTARG" >&2; usage >&2; exit 2 ;; :) printf '%s: option -%s requires a value\n' "${0##*/}" "$OPTARG" >&2; exit 2 ;; ``` For a deploy script the stakes are the point: a flag the user believed was a dry-run switch, quietly discarded, means a real deployment. ## Bug two: help is not an error `-h` is the user asking for something and getting it. That is a success: - text on **stdout**, so `./deploy.sh -h | less`, `> help.txt` and `grep` all work; - exit status **0**, so a CI step, a `set -e` wrapper or a smoke test that runs the tool with `-h` does not fail. A usage *error* is the opposite: - message and usage text on **stderr**, so it does not contaminate a pipeline's data stream; - a **non-zero** exit status. A script that writes help to stderr, or exits non-zero from its `-h` branch, breaks every automated caller — which is exactly the CI symptom described. ## Why 2 for a usage error There is no standard that mandates it, but a widely followed convention reserves a distinct status for "you called me wrong", separate from "I ran and the work failed". `grep` uses 1 for "no lines matched" and 2 for an error; `diff` uses 1 for "files differ" and 2 for trouble. Following that, `exit 2` for a bad command line and `exit 1` for a genuine failure lets a caller tell a broken invocation from a broken deployment without parsing text. What matters most is that the choice is documented in the script's header and used consistently across a team's tools. ## Writing usage once Have one function that emits the text on stdout, and let each call site decide the stream and the status. That keeps help and errors from drifting apart: ```bash usage() { cat <<EOF Usage: ${0##*/} [-v] [-f FILE] TARGET -f FILE write output to FILE -v verbose -h show this help EOF } ``` A here-document keeps the text readable, and `${0##*/}` prints the script's base name so the message matches whatever the user typed. Resist making `usage` exit on its own: a function that always exits 2 cannot serve the `-h` path, and one that always exits 0 cannot serve the error path. ## Finish the validation Option parsing is only half of the contract. After `shift $((OPTIND - 1))`, check the operands too — wrong count, a missing required argument, an operand that starts with a dash because the user put a flag after it — and route every one of those through the same stderr-plus-exit-2 path. A single, predictable failure mode is what makes a script safe to wrap. ## What good looks like - Silent-mode optstring, with `\?)` and `:)` branches that exit. - `-h` prints usage on stdout and exits 0. - Every usage error prints to stderr and exits 2. - Operand count validated after the shift, through the same path. - The exit-status meanings written in a comment at the top of the file.
- Why does `set -euo pipefail` not catch the mistyped flag?Because nothing fails. getopts returns success when it reports an invalid option, so `set -e` has no non-zero status to act on, and every variable involved was already initialised, so `set -u` has nothing unset to trap. The guard has to be an explicit `\?` branch that exits.
- Should the usage function itself call exit?No — keep it a pure emitter that writes to stdout, and let each caller choose. The `-h` branch runs `usage; exit 0`; an error branch runs `usage >&2; exit 2`. A usage function with a baked-in exit status can only serve one of the two paths correctly.
- What should the script do when the flags parse cleanly but the wrong number of operands is given?Treat it as the same class of failure: message and usage on stderr, exit 2. Validate right after `shift $((OPTIND - 1))`, for example checking the argument count and rejecting any remaining operand that begins with a dash. Callers can then rely on one status meaning "bad command line".
saying these in an interview costs you the question
- Believes a bad option ends the getopts loop
- Expects set -e to catch an unrecognised flag
- Prints the help text to stderr
- Exits non-zero after a successful -h request
- Uses the same exit status for usage errors and runtime failures