skip to content

Process and Job Control

What actually forks, what inherits what, and how a script cleans up after itself. Subshells, command substitution, background jobs, traps, and exec are the mechanics behind "my variable disappeared" and "my temp files were left behind" — two bugs interviewers love to reach for.

part ofBashoverview, primer and where to startread it →
on this pageshow

questions

21

In a bash script, what does appending & to a command do, what does the special parameter $! hold afterwards, and what does a bare wait do?

level: juniorimportance: must knowfreq 68%

answer

  1. runs it, doesn't stop for it
  2. status is 0 before anything finished
  3. a one-slot register holding a PID
  4. overwritten by the next background job
  5. blocking until the children are done

basics

~20 s

Appending & runs the command asynchronously in a background child process and returns immediately with status 0. $! then holds that child's PID. A bare wait blocks until every background child of the script has finished.

solid answer

~40 s

`cmd &` starts `cmd` in a forked child and does **not** wait for it — the shell moves straight to the next line, and the status of the asynchronous command itself is always 0, so you learn nothing yet about whether `cmd` worked. Immediately afterwards `$!` expands to the PID of that child, and it is overwritten by the next `&`, so you capture it right away, usually into an array. `wait` with no arguments blocks until all of the script's currently running background children have terminated; `wait "$pid"` blocks for one child and returns *its* exit status. A script that backgrounds work and then falls off the end without waiting simply exits while the children are still running, which in CI usually means truncated output or work killed mid-flight.

code

bash · 13 lines
bash
#!/usr/bin/env bash
# Launch two jobs, capture each PID at once, then collect them.
sleep 2 &
pid_a=$!
sleep 1 &
pid_b=$!

echo "started $pid_a and $pid_b"

wait "$pid_a"
echo "a finished with $?"
wait "$pid_b"
echo "b finished with $?"

go deeper

for a junior

Be able to say plainly that & runs a command in the background without waiting, that $! is the PID of that job, and that wait blocks until background children finish. Show the three-line capture idiom.

for a middle

Explain that the asynchronous command's own status is always 0, that $! is overwritten by the next &, and that for a backgrounded pipeline $! names the last command in it. Know that wait only works on children of this shell.

for a senior

Show why a script that exits without waiting produces truncated CI logs and work killed mid-flight, and treat PID capture into an array as the default shape of any fan-out script you would put in production.

for a principal

Frame & as introducing supervision debt: every forked child needs an owner that collects it, a bound on how many exist, and a defined teardown path. Be ready to argue when a shell script should not be the thing doing the supervising at all.

## What `&` actually does In bash, `&` is a *command terminator*, exactly like `;` — it ends the command in front of it. The difference is that `;` runs the command and waits for it, while `&` runs it **asynchronously**: bash forks a child, the child runs the command, and the parent shell returns immediately to read the next line. Two consequences follow directly and both surprise people: 1. **The exit status you see is not the command's.** An asynchronous command's status is 0 the moment it is launched, because nothing has finished yet. `cmd & echo $?` prints 0 even if `cmd` is going to fail two seconds later. 2. **The command runs in its own process**, so anything it changes about the shell — variables, the current directory — is not visible to the parent afterwards. (The mechanics of that isolation are the subshell topic; here it is enough to know a background job is a separate process.) In a non-interactive script, job control is off by default, so bash does not print the `[1] 12345` job line you see at an interactive prompt. The job still exists; `jobs` and `jobs -p` list it from within the same shell. ## `$!` — the PID of the last background job After `&`, the special parameter `$!` expands to the process ID of the job just placed in the background. It is a single value that is **overwritten by the next `&`**, which is why the idiom is always to read it on the same line or the very next one: ```bash long_task a & pid_a=$! long_task b & pid_b=$! ``` If you background two jobs and only then read `$!`, you have the second PID and have permanently lost the first — there is no way to recover it later. For fan-out, collect into an array: `pids+=("$!")`. For a backgrounded **pipeline** such as `producer | consumer &`, `$!` is the PID of the *last* command in the pipeline (`consumer`). The whole pipeline is one job, but `$!` names only that one process. ## `wait` `wait` is a shell builtin — it can only wait for children of *this* shell, which is why you cannot `wait` for an arbitrary PID on the machine (you get status 127 instead). - `wait` with no arguments: block until every currently running background child has terminated. - `wait "$pid"`: block until that one child terminates, and return **its** exit status. - If the child was killed by a signal, the status is 128 plus the signal number (137 for SIGKILL). ```bash sleep 5 & pid=$! wait "$pid" echo "sleep finished with $?" ``` ## Why a script needs `wait` at all If the script reaches its last line while children are still running, bash exits anyway. The children do not die with it — they keep running, now with no parent shell tracking them. That produces the classic CI symptom: the job "passes" in ten seconds because the real work was still in flight when the script returned, and its output never made it into the log. Worse, whatever supervises the script (a CI runner, a container entrypoint, a systemd unit) may tear down the whole process group a moment later and kill the work half-done. So the shape of every fan-out script is: start the children, remember their PIDs, and `wait` before doing anything that depends on their results — including exiting. ## Related things `&` is *not* - `&` is not `&&`. `a & b` runs `a` in the background and then runs `b` immediately; `a && b` runs `b` only if `a` succeeded. - `&` does not make anything faster on its own. It overlaps waiting time; if the work is CPU-bound and you background more jobs than you have cores, you mostly add contention. - `&` alone does not detach a job from the terminal or from the script's process group — `nohup` and `disown` address that, and they are a different question. ## The mental model Treat `&` as "fork and forget", `$!` as a one-slot register holding the last fork's PID, and `wait` as the only way to turn that fork back into a result. A script that uses `&` without `$!` and `wait` has started work it cannot supervise.

  • If you background a pipeline with `producer | consumer &`, whose PID ends up in `$!`?
    The last command of the pipeline — `consumer`. Bash treats the pipeline as a single job, but `$!` names only its final process, and `wait "$!"` therefore returns that command's status, not the producer's. If you need the producer's outcome too, that is what `pipefail` and `PIPESTATUS` are for.
  • Why must `$!` be read immediately after the `&` rather than later in the script?
    `$!` holds only the most recently backgrounded job's PID and is overwritten by the next `&`. Once you background a second job, the first PID is gone for good — there is no history. The safe idiom is to assign it on the spot, typically appending to an array: `cmd & pids+=("$!")`.
  • What happens if a script exits while its background children are still running?
    Bash exits without killing them; the children keep running, now orphaned from the script that started them. Their output may never reach the log, and whatever supervises the script may tear down the process group and kill them mid-work. If you care about the result, `wait` for them; if you do not, kill them explicitly.

saying these in an interview costs you the question

  • Says $! holds the background job's exit status
  • Thinks & by itself makes a command run faster
  • Reads $! after launching a second background job
  • Believes the script blocks until background work finishes
  • Confuses & with && in a command chain

context

open as a page

In a bash script, what is the difference between writing a command substitution as `$(cmd)` and as the older backtick form `` `cmd` ``, and why does nesting one inside another behave differently?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Both run a command and substitute its standard output, but $(...) nests and quotes cleanly because everything up to the matching ) is parsed as a fresh command, while backticks need a backslash escape for every nested level and mangle backslashes and inner quotes.

open as a page

A bash script sets GREETING=hello and then runs ./child.sh, but child.sh prints an empty value for $GREETING. Why doesn't the child see it, and what makes a variable visible to child processes?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Only exported variables become part of a child process's environment. A plain assignment creates a shell variable that lives in the current shell alone; marking it with export copies it into every command the shell runs afterwards, including child.sh.

open as a page

In a bash script, `trap cleanup EXIT` is the usual way to guarantee cleanup. Which ways of the script ending actually run that EXIT trap, and which ones do not?

level: middleimportance: must knowfreq 62%

basics

~20 s

Bash runs an EXIT trap whenever the shell itself decides to exit: the end of the script, any explicit exit, an abort under set -e, and after a fatal signal such as SIGTERM. SIGKILL leaves it unrun.

open as a page

In a bash script, `count=0; cat ids.txt | while read -r id; do count=$((count+1)); done; echo "$count"` prints 0 however many lines the file has. Why does the count vanish, and how do you make it survive the loop?

level: middleimportance: must knowfreq 75%

basics

~20 s

Every stage of a bash pipeline runs in its own forked subshell, so the while loop increments a copy of count that dies when the pipeline finishes. Feed the loop by redirection instead of through a pipe.

open as a page

A container's entrypoint is a bash script whose last line is "$@". Stopping the container always takes the full ten-second timeout and the application never logs its shutdown message. Explain what the shell is doing, and what change to that last line fixes it.

level: seniorimportance: must knowfreq 58%

basics

~20 s

The bash script is the container's main process, and the application is only its child. The stop signal goes to bash, which neither forwards it nor acts on it, so the application is eventually killed outright. Changing the last line to exec "$@" makes the application the main process itself.

open as a page

In bash, what is the difference between grouping commands with `( ... )` and with `{ ...; }`, and what happens to a `cd` or a `umask` performed inside each form?

level: juniorimportance: should knowfreq 60%

basics

~20 s

Parentheses run the group in a forked subshell; braces run it in the current shell. A cd, umask or variable assignment inside parentheses is undone when the group ends, while the same thing inside braces changes the running shell.

open as a page

A bash backup script starts three uploads with & and then calls a bare wait, yet it exits 0 even when one upload fails. Why does the bare wait hide the failure, and how do you collect each child's exit status?

level: middleimportance: should knowfreq 52%

basics

~20 s

A bare wait returns 0 once all children are collected — it never reports their statuses. To detect failure, save each PID from $! into an array and call wait on each PID separately, since wait with a PID argument returns that child's exit status.

open as a page

A bash script does `body=$(cat report.txt)` on a file that ends with a blank line, and a later `printf '%s' "$body" | wc -c` reports fewer bytes than the file. What does command substitution do to the captured output, and how would you capture it byte-for-byte?

level: middleimportance: should knowfreq 50%

basics

~20 s

Command substitution strips every trailing newline from the captured output, not just the last one, so a file ending in blank lines loses them. To keep them, append a sentinel inside the substitution and remove it afterwards: body=$(cat report.txt; printf x); body=${body%x}.

open as a page

In bash, what does the leading assignment in the command `LC_ALL=C sort names.txt` do, and how is it different from putting `LC_ALL=C` on its own line before the sort?

level: middleimportance: should knowfreq 44%

basics

~20 s

An assignment written in front of a command puts that name into the environment of that one command only, leaving the shell's own variables untouched. On its own line it creates an ordinary shell variable, which is not exported, so sort never sees it at all.

open as a page

In a bash script, what is the difference between the line `exec >>/var/log/job.log 2>&1` and the line `exec /usr/bin/myjob`?

level: middleimportance: should knowfreq 48%

basics

~20 s

With a command, exec replaces the running shell with that program: same process, same PID, and nothing after the line ever runs. With only redirections and no command, exec applies those redirections to the current shell permanently, and the script continues normally.

open as a page

In a bash script, what is the difference between `trap 'handler' TERM`, `trap '' TERM` and `trap - TERM`, and which of the three still applies to commands the script launches?

level: middleimportance: should knowfreq 42%

basics

~20 s

A non-empty argument sets a handler, an empty string makes the shell ignore the signal, and a lone dash restores the default action. Only the ignore is inherited by commands the script launches; handlers are not.

open as a page

In a bash script, `echo $$` inside a `( ... )` subshell prints the same number as `echo $$` outside it. Why doesn't `$$` change, and which variable reports the PID of the process actually running that code?

level: middleimportance: should knowfreq 35%

basics

~20 s

Bash keeps $$ pinned to the process ID of the shell that was invoked, even inside a subshell, because POSIX defines it that way. Use $BASHPID, added in bash 4.0, to get the PID of the process currently executing.

open as a page

A bash script must run the same task over 500 inputs but never have more than 8 running at once. How do you bound the concurrency, and how do you still find out which inputs failed?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Two idioms bound it: pipe the inputs into xargs -P 8 -n1, or keep a PID list in bash and call wait -n whenever 8 jobs are in flight. Neither reports which input failed unless the worker records failures itself.

open as a page

This bash script runs under `set -euo pipefail`, yet it prints `deploying` with an empty version and carries on after `get_version` fails: get_version() { cat VERSION; } main() { local ver=$(get_version); echo "deploying $ver"; } Why does the failure not stop the script, and what is the fix?

level: seniorimportance: should knowfreq 46%

basics

~20 s

local is itself a command, and its exit status — not the substitution's — becomes the status of the line, so the failure is masked and set -e never fires. Declare and assign on separate lines: local ver; ver=$(get_version). ShellCheck flags the original as SC2155.

open as a page

A bash entrypoint script sets `trap shutdown TERM` and then runs a long-lived server as its last foreground command. When the supervisor sends SIGTERM, nothing happens until the server exits on its own. Why is the handler delayed, and how do you restructure the script so it runs at once and forwards the signal to the server?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Bash finishes the foreground command before running a trap handler, so a trap set around a blocking command fires only after it returns. Run the server in the background instead and block in wait, which a trapped signal interrupts immediately.

open as a page

In a bash script, what does `kill -0 "$pid"` do, and what does its exit status tell you?

level: juniorimportance: nice to knowfreq 30%

basics

~20 s

kill -0 sends no signal at all. It performs only the existence and permission check that kill would do first, exiting 0 when a process with that PID exists and you are allowed to signal it, and non-zero otherwise.

open as a page

In bash, what does the `$(<file)` form of command substitution do differently from `$(cat file)`, and when would you not use it?

level: middleimportance: nice to knowfreq 24%

basics

~20 s

$(<file) makes bash read the file's contents itself instead of forking and running an external cat, so it is faster and substitutes the same text. It is a bash and ksh extension rather than POSIX, so a strict /bin/sh script keeps $(cat file).

open as a page

A long-running bash script installs a newer build of a tool as /usr/local/bin/mytool, but later calls to mytool in the same script still run the old /usr/bin/mytool. How does bash resolve a bare command name, and what makes it keep using the stale path?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

Bash remembers the full path of each external command it has run, in a per-shell hash table, and reuses it instead of searching PATH again. The old location is still valid, so the cached entry keeps winning. Clearing it with hash -r forces a fresh lookup.

open as a page

In bash, what does `shopt -s lastpipe` change about how a pipeline is executed, what condition must hold for it to take effect, and when would you rely on it in a production script?

level: seniorimportance: nice to knowfreq 25%

basics

~20 s

With lastpipe enabled, bash runs the final command of a pipeline in the current shell instead of a forked subshell, so its variable assignments survive. It only applies when job control is inactive, which is the default in scripts but not at an interactive prompt.

open as a page

A bash deploy script fans work out to background jobs on a shared CI runner. How do you decide the concurrency limit, and would you abort the remaining jobs at the first failure or let everything run and aggregate the failures?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Set the limit from the bottleneck resource — cores for CPU work, the remote rate limit or connection pool for network work, memory per child for everything — not from a core count the container may not actually own. Choose fail-fast for mutating work and aggregate reporting for independent checks.

open as a page