A bash backup script starts three uploads with & and then calls a bare wait, yet it exits 0 even when one upload fails. Why does the bare wait hide the failure, and how do you collect each child's exit status?
answer
- a barrier, not an error check
- its status is always zero
- strict mode never sees the child fail
- one PID at a time returns the real status
- above 128 means a signal, not an exit code
basics
~20 sA bare wait returns 0 once all children are collected — it never reports their statuses. To detect failure, save each PID from $! into an array and call wait on each PID separately, since wait with a PID argument returns that child's exit status.
solid answer
~50 s`wait` with no arguments waits for every background child and then returns 0; it discards the individual statuses, so the script cannot tell a clean run from a failed upload. `set -e` does not help either — a background job failing is not an errexit trigger. The fix is to keep the PIDs: append `$!` to an array as each job starts, then loop over the array calling `wait "$pid"`, because that form returns the child's own status. Record failures as you go — `wait "$pid" || rc=1` — and exit non-zero at the end, or map the PID back to the input so the log says *which* upload failed. A status above 128 means the child was killed by a signal (137 is SIGKILL); 127 means the PID was not a child of this shell.
code
bash · 19 lines#!/usr/bin/env bash
set -euo pipefail
targets=(alpha beta gamma)
pids=()
for t in "${targets[@]}"; do
( sleep 1; [[ $t != beta ]] ) & # 'beta' fails
pids+=("$!")
done
rc=0
for i in "${!pids[@]}"; do
if ! wait "${pids[i]}"; then
echo "job ${targets[i]} failed with status $?" >&2
rc=1
fi
done
exit "$rc"go deeper
Know that starting a job with & gives you no result, and that a bare wait only tells you everything finished. Remember to save $! if you will ever need that job's outcome.
Explain that wait with no argument returns 0 while wait with a PID returns that child's status, and write the PID-array collection loop from memory, including accumulating a non-zero exit code.
Be able to say why strict mode gives false confidence here, read 137 and 143 as signal terminations rather than application errors, and insist that a fan-out script names which input failed rather than reporting a bare non-zero exit.
Own the policy question: whether the script aborts on the first failure or reports a complete list, how partial success is surfaced to the caller, and at what point per-job supervision belongs in a real orchestrator instead of a wait loop.
## The bug ```bash upload a & upload b & upload c & wait echo "all done" ``` This script is honest about only one thing: all three children have terminated. It says nothing about *how* they terminated, and its own exit status is 0 regardless. There are two separate reasons the failure is invisible: 1. **A bare `wait` returns 0.** When `wait` is given no argument it waits for all currently active children and its return status is zero. Individual statuses are collected by the shell and thrown away. 2. **`set -e` does not fire for background jobs.** Errexit reacts to the failure of commands the shell runs and waits for. Launching `upload a &` succeeds (status 0) the moment the fork happens; the child's later failure is not a command failure in the parent. So even a strict-mode script sails past it. ## What `wait` returns in each form | Form | Returns | |---|---| | `wait` | 0, after all background children terminate | | `wait "$pid"` | that child's exit status | | `wait -n` | the status of the *next* child to terminate (bash 4.3+) | Two status values are worth memorising because they show up in logs: - **128 + N** means the child was terminated by signal N. 137 = 128 + 9 (SIGKILL, typically the OOM killer or a runner timeout); 143 = 128 + 15 (SIGTERM). - **127** from `wait "$pid"` means the shell has no such child — usually because the PID was already reaped, or you passed a PID from a different shell. `wait` is a builtin and can only collect this shell's own children. ## The collection pattern The standard shape keeps a PID array and, when the message matters, a parallel array of labels: ```bash pids=() names=() for target in a b c; do upload "$target" & pids+=("$!") names+=("$target") done rc=0 for i in "${!pids[@]}"; do if ! wait "${pids[i]}"; then echo "upload ${names[i]} failed with $?" >&2 rc=1 fi done exit "$rc" ``` Points that make this work rather than merely look right: - `$!` is captured **inside** the loop, immediately after the `&`. Read it later and you have only the last PID. - The status is checked per PID, so the script knows which target failed, not just that something did. - `rc` accumulates instead of returning early, so one bad upload does not hide a second one. - Under `set -e`, a bare `wait "$pid"` that returns non-zero *would* abort the script — which is sometimes what you want, but inside `if ! wait ...` it is a condition, so errexit is suspended and you keep control. A common variant collects into an associative array keyed by PID, `declare -A job_name; job_name[$!]="$target"`, which is tidier but needs bash 4+. ## `wait -n` and its version fence `wait -n` (bash 4.3 and newer) returns as soon as *any* one child finishes, giving you its status — the building block for fail-fast fan-out and for keeping a fixed number of jobs in flight. Its weakness is that plain `wait -n` tells you the status but not *which* job it belongs to; bash 5.1 added `wait -n -p varname`, which stores the PID of the job that finished into `varname`. Version matters here in practice: macOS still ships bash 3.2 as `/bin/bash`, which has neither `wait -n` nor associative arrays. A script that must run there uses the PID-array loop above, which works everywhere. ## Related traps - **Killed vs failed.** If your supervisor kills the whole process group, every child comes back as 143 or 137. Treat >128 as "terminated by a signal", not as an application error code. - **Reaping order.** The loop waits for PIDs in start order, so a fast third job is collected last. That is fine for correctness — bash keeps the status until you ask — but it means the script's total runtime is the slowest job, and errors are reported in start order rather than failure order. - **Output interleaving.** Three children writing to the same stdout interleave arbitrarily. If the log has to be readable, give each child its own output file and print them after the wait loop. ## The one-line summary `wait` alone is a barrier, not an error check. Anything that cares whether the background work succeeded has to wait on PIDs one at a time and aggregate the statuses itself.
- Does `set -euo pipefail` at the top of the script change any of this?No. Launching an asynchronous command succeeds immediately, so errexit has nothing to react to when the child later fails, and a bare `wait` returns 0 so it does not trip either. Strict mode only helps once you call `wait "$pid"` as a plain command — then a non-zero status does abort the script.
- `wait "$pid"` returns 137. What does that tell you?The child was terminated by signal 9 (SIGKILL), since a signalled child reports 128 plus the signal number. It did not choose to exit with 137. In practice that usually means the OOM killer, a CI job timeout, or an explicit `kill -9` — so you investigate the environment, not the program's error handling.
- What does `wait -n` add, and what can it not tell you on its own?`wait -n` (bash 4.3+) returns as soon as any one background job finishes and gives you that job's exit status, which is what fail-fast and fixed-concurrency loops need. On its own it does not say *which* job finished; bash 5.1 added `wait -n -p varname` to store the finished PID, and before that you compare against your PID list.
saying these in an interview costs you the question
- Believes a bare wait returns the failing child's status
- Says set -e catches a failing background job
- Reads $? after wait instead of waiting per PID
- Treats exit status 137 as an application error code
- Assumes wait can collect any PID on the machine