skip to content

A bash deploy script calls `run_migration`, then finishes with `echo "migration done"`, and the CI job reports success even when the migration fails. Why does the script exit 0, and how do you make it report the real status?

level: seniorimportance: nice to knowfreq 34%

answer

  1. the last command decides the script's status
  2. a log line succeeds and hides the failure
  3. bare exit means exit with the current status
  4. capture it on the very next line
  5. force a failure and print the status

basics

~20 s

A bash script's exit status is the status of the last command it ran, and here that is the successful echo, which masks the migration failure. Capture the status immediately with rc=$? and end with exit "$rc", or fail early on the migration itself.

solid answer

~50 s

When a script ends without an explicit `exit`, bash reports the exit status of the last command it executed — and a trailing `echo` almost always succeeds, so it overwrites whatever the migration returned. A bare `exit` with no argument has the same problem in reverse: it means `exit $?`, so it faithfully propagates whatever ran immediately before it, which may well be a log line rather than the real work. The fixes are all about making the status you care about the one that reaches the caller: capture it on the very next line with `rc=$?` and finish with `exit "$rc"`, or fail fast at the point of failure with `run_migration || exit 1`, or wrap it in `if ! run_migration; then ...; exit 1; fi`. This matters because CI, cron, container orchestrators and monitoring all decide pass or fail from that single number.

code

bash · 7 lines
bash
#!/usr/bin/env bash
run_migration() { echo "migration exploded" >&2; return 1; }

run_migration
rc=$?
echo "migration finished (status $rc)"
exit "$rc"

go deeper

for a junior

Remember that a script's exit status is whatever its last command returned, and that a successful echo at the bottom hides an earlier failure.

for a middle

Explain both shapes of the bug — running off the end, and a bare exit that means exit $? — and write the capture-and-propagate fix with rc=$? on the immediately following line.

for a senior

Show the operational instinct: verify propagation by forcing a failure and printing $?, read the bottom of every script in review, and preserve the original code when callers branch on it.

for a principal

Own it as a reliability property. A permanently green job is worse than a red one, so treat exit-status propagation as part of the definition of done for operational scripts and make forced-failure verification a standard acceptance step.

## The rule that causes it A bash script's own exit status is: - the argument to `exit`, if the script calls `exit N`; or - the status of the **last command executed**, if the script simply runs off the end; and - a bare `exit` with no argument is defined as `exit $?` — the status of whatever ran immediately before it. So the last thing your script does is what the outside world sees. If that last thing is a friendly progress message, the outside world sees success. ```bash #!/usr/bin/env bash run_migration # returns 1 echo "migration done" # returns 0 # script ends here -> exit status 0 ``` The symptom is a CI job that is permanently green, a cron job that never mails a failure, a container that a scheduler considers to have completed cleanly. It is one of the most consequential shell bugs precisely because nothing looks wrong: the log even contains the migration's error output, but the number that the platform gates on says success. ## Where it hides The pattern shows up in several disguises: - **A trailing summary line.** `echo "finished at $(date)"` at the bottom of every script is a house style in many teams and silently masks the status. - **A bare `exit` after a cleanup step.** `cleanup; exit` propagates `cleanup`'s status, not the work's. - **A final conditional.** If a script ends with `if [ -n "$notify" ]; then send_mail; fi` and the condition is false, the `if` produces `0` — a false branch is not a failure — and that becomes the script's status. - **A loop or a redirect at the end.** Any last construct produces a status of its own. ## The fixes, in order of preference **1. Fail at the point of failure.** The best scripts never reach the end in a broken state: ```bash run_migration || { echo "migration failed" >&2; exit 1; } echo "migration done" ``` or equivalently with a conditional: ```bash if ! run_migration; then echo "migration failed" >&2 exit 1 fi ``` Now the trailing `echo` is only ever reached on the success path, so it cannot mask anything. **2. Capture and propagate.** When you must run something after the failure — a report, an upload of logs — save the status before anything else executes: ```bash run_migration rc=$? upload_logs echo "migration finished (status $rc)" exit "$rc" ``` The capture must be the *immediately* following line; anything in between overwrites the value, which is the same one-shot property that makes `$?` easy to misuse. **3. Preserve the distinction you care about.** If a caller branches on specific codes, propagate the original number rather than collapsing everything to `1`. `exit "$rc"` keeps a tool's own `2` or `3`; `exit 1` throws that information away. Conversely, if you want a stable contract, map the codes deliberately and document the mapping. ## Related traps in the same family A cleanup handler registered to run on exit can also overwrite the status if its last command succeeds — the standard guard is to capture the status at the top of the handler and re-exit with it; the mechanics of registering such a handler belong with signals and traps. Bash's automatic abort-on-error setting is a separate lever with its own well-known blind spots, covered under strict mode and error handling; note only that it does not save you here, because the trailing `echo` genuinely succeeds and there is nothing for it to abort on. ## How to catch it in review Read the bottom of every script and ask: *what is the last command on the success path, and on each failure path?* If the answer on any failure path is a log line, the script lies to its caller. A quick empirical check is worth more than reasoning: ```bash ./deploy.sh; echo "exit status: $?" ``` Run it once with the real work forced to fail. If it still prints `0`, you have found the bug. Making that check part of how you accept a new operational script is cheap and catches the whole class. ## What to say in an interview "The script exits with the status of its last command, which is the `echo`. I'd either fail fast with `run_migration || exit 1`, or capture `rc=$?` on the next line and end with `exit "$rc"` — and I'd verify by running the script with the step forced to fail and printing `$?`."

  • What status does a bash script exit with if it never calls `exit` at all?
    The status of the last command it executed. That is why a trailing `echo`, a final `if` whose condition was false, or a cleanup call at the bottom all become the script's public result. Running off the end is not an implicit success — it is an implicit inheritance from whatever happened to run last.
  • Why is `exit 1` for everything sometimes the wrong fix?
    It collapses information. If a caller branches on distinct codes — bad usage versus a missing dependency versus a failed remote call — flattening them all to 1 removes that ability, and a tool's own meaningful 2 or 3 is lost. Propagate the captured status with `exit "$rc"` when the codes matter, and reserve a deliberate mapping for when you want a stable contract.
  • How would you prove a script propagates failure correctly before shipping it?
    Run it with the real work forced to fail — point it at a bad host, stub the function to return non-zero — and then print `$?` immediately after. If it prints 0, the script lies to CI. Making this a required check when accepting an operational script catches the whole class cheaply, since the bug is invisible on the happy path.

saying these in an interview costs you the question

  • Assumes a script that finishes normally always exits 0
  • Thinks a bare exit defaults to 0
  • Adds a trailing log line without considering the status
  • Believes the loudest error in the log determines the exit code
  • Collapses every failure to exit 1 and loses the caller's distinctions

context