skip to content

Scripting Best Practices

The difference between a snippet that works on your laptop and a script you are willing to run from cron on a production host: strict mode, real error handling, linting, deliberate portability choices, and predictable CLI behavior. Senior shell interviews spend most of their time in this area.

part ofBashoverview, primer and where to startread it →
on this pageshow

questions

page 1 of 2

A bash script fails only on the CI machine and you cannot reproduce it locally. What does `set -x` add to the script's output, where does that output go, and how do you enable it for one section only — or for a single run without editing the file at all?

level: juniorimportance: must knowfreq 74%

answer

  1. shows what ran, not what you wrote
  2. expanded values, before execution
  3. goes to stderr, not stdout
  4. set +x turns it back off
  5. bash -x needs no file edit

basics

~20 s

set -x makes bash print every command to stderr just before running it, after expansion, prefixed by PS4 (default "+ "). Wrap a region in set -x and set +x to scope it, or run bash -x script.sh to trace one run without touching the file.

solid answer

~50 s

`set -x` turns on xtrace: before bash runs each simple command it writes that command to **stderr**, prefixed by the value of `PS4` (default `+ `). The key point is that the trace is printed **after expansion** — you see `+ rm -rf /var/tmp/build 42` with the variables, globs and command substitutions already resolved, which is exactly the gap between what the script says and what it does. Because it goes to stderr, `./script.sh > out.log` does not capture it and `2>/dev/null` silently throws it away. To scope it, put `set -x` before the suspicious region and `set +x` after it — the setting is shell-global, not block- or function-scoped, so if a helper turns it off it stays off for the caller too. To trace a run without editing anything, invoke the script as `bash -x ./script.sh`. `set -v` is the sibling option: it echoes each input line as bash *reads* it, before expansion.

go deeper

for a junior

Know that set -x prints each command before bash runs it, that the output is on stderr, and that set +x turns it off again. Be able to say you would run bash -x ./script.sh to trace a failing run.

for a middle

Explain that the trace is emitted after expansion, so it reveals word splitting, glob results and substituted values, and that the PS4 prefix's repeated first character marks nesting. Show that you know options are shell-global, not block-scoped.

for a senior

Talk about operating it: gating tracing behind a DEBUG environment variable, redirecting stderr so the trace does not corrupt machine-parsed stdout, the log volume a traced loop produces, and the fact that expanded commands print credentials in clear.

for a principal

Frame it as a policy question — what every script your organisation ships should support: a documented switch to turn tracing on for one run, a rule that traces never land in a shared log unredacted, and where structured logging from the script beats a raw shell trace.

## The problem xtrace solves A shell script is a program whose text and its runtime meaning can differ enormously, because every line goes through quote removal, parameter expansion, command substitution, word splitting and pathname expansion before a command is executed. Reading the source tells you what you *wrote*; `set -x` tells you what bash actually *ran*. ## What it prints, and when With `set -x` (equivalently `set -o xtrace`) in effect, bash writes each simple command, `for`/`case`/`select` word list and arithmetic `for` expression to standard error immediately before executing it. The line is preceded by the expanded value of the shell variable `PS4`, whose default value is `+ `. The critical property is *when* it prints: after expansion, before execution. Given ```bash file="my report.txt" set -x rm $file ``` the trace reads `+ rm my report.txt`, and you can immediately see that `rm` received two arguments because `$file` was unquoted. Nothing about the source line hints at that; the trace does. The first character of `PS4` is repeated to show nesting: a command running one level deeper — inside a command substitution or a subshell — is prefixed `++` rather than `+`. That is how you tell a command in `$( )` apart from one at the top level of the script. ## Where the output goes xtrace is written to **stderr**, not stdout. Three consequences bite people constantly: - `./deploy.sh > deploy.log` captures the script's normal output and leaves the trace on the terminal. - `./deploy.sh 2>/dev/null` throws the whole trace away. - If the script's own diagnostics also go to stderr, the two interleave, which is what makes the trace hard to read in a CI log. To capture everything in order, redirect both: `bash -x ./deploy.sh > run.log 2>&1`. ## Turning it on and off `set -x` enables it, `set +x` disables it — the `+` form of a `set` flag always turns the option off. Shell options are properties of the shell, not of a block, so enabling xtrace inside a function leaves it enabled after the function returns. If you want a helper to leave the caller's setting untouched, save and restore it yourself. The special parameter `$-` holds the currently enabled option flags: ```bash noisy_step() { local had_x=$- set -x tar -czf "$out" "$dir" case $had_x in *x*) ;; *) set +x ;; esac } ``` ## Tracing without editing the file Often you cannot or should not edit the script — it is baked into a container image, or owned by another team. Two ways in: - Invoke the interpreter explicitly: `bash -x ./script.sh args...`. This runs the whole script traced. - Export `SHELLOPTS` with `xtrace` in it. `SHELLOPTS` is a read-only, colon-separated list of enabled `set` options; if it is present in the environment when bash starts, those options are enabled before anything else runs, so `export SHELLOPTS=xtrace` propagates tracing into child bash scripts too. That is a blunt instrument — it traces *everything*, including scripts you did not mean to trace. ## `set -v`, the other half `set -v` (verbose) echoes each line of input as the shell reads it — the raw source, before any expansion. Used alone it is rarely what you want; used together (`bash -xv`) you see each source line followed by the expanded commands it produced, which is useful when a line expands into something unrecognisable. ## Costs and cautions xtrace is not free and not neutral: - **Volume.** A loop over a thousand files produces thousands of trace lines, which can dominate a CI log or fill a disk. - **Secrets.** Because commands are printed *expanded*, a token in a variable is printed in clear. A script that traces `curl -H "Authorization: Bearer $TOKEN"` puts the token in the log. - **Ordering.** Trace goes to stderr while the program's output usually goes to stdout; when they are redirected to different places, or buffered differently, the apparent ordering can mislead. So the normal pattern is not "always on": gate it behind a flag — `[[ ${DEBUG:-0} == 1 ]] && set -x` — so an operator can turn on tracing for one run without a code change, and the default run stays quiet.

  • Why does redirecting a traced script's stdout to a file leave the trace on your terminal?
    Because xtrace is written to standard error, not standard output. `> run.log` only redirects fd 1, so the trace keeps going to fd 2. Capture both with `> run.log 2>&1`, or route the trace to its own descriptor so the two streams never interleave.
  • A helper function calls `set +x` at the end. Why does the rest of the script stop tracing too?
    Shell options are global to the shell, not scoped to a function or block, so `set +x` inside a function disables xtrace for everything that runs after it returns. A well-behaved helper saves the incoming state — `local had_x=$-` — and only turns tracing off again if it was off when it was entered.
  • How would you trace a script that is baked into a container image and that you cannot modify?
    Override the entrypoint to run the interpreter explicitly: `bash -x /app/entrypoint.sh`. If the script is invoked indirectly by something else, `export SHELLOPTS=xtrace` in the container environment enables xtrace in every bash that starts, because bash reads that variable at startup. The second option is blunt and traces unrelated scripts too.

saying these in an interview costs you the question

  • Thinks set -x prints source lines exactly as written
  • Expects the trace on stdout, so > log captures it
  • Confuses set -x with set -e or set -v
  • Believes set -x inside a function is scoped to it
  • Leaves tracing on permanently, printing tokens into logs

context

open as a page

A bash deploy script assembles a command string from an environment variable it did not set and runs it with eval "$cmd". What does eval do to that string, and what can whoever controls the variable achieve?

level: juniorimportance: must knowfreq 68%

basics

~20 s

eval joins its arguments and hands the result back to the shell parser, so metacharacters inside the variable become syntax instead of data. Anyone who controls the value can run arbitrary commands as the script's user.

open as a page

A bash script parses flags with `while getopts "o:v" opt; do ... done`, then reads its input file from `$1` — but when it is run as `script -v data.csv`, `$1` is `-v` instead of `data.csv`. What is OPTIND, and which line is missing after the loop?

level: middleimportance: must knowfreq 62%

basics

~20 s

OPTIND is the index of the next argument getopts will examine, and getopts never alters the positional parameters itself. The missing line is shift $((OPTIND - 1)) after the loop, which drops the parsed options so operands start at $1.

open as a page

One nightly cron job stopped working weeks ago and nobody noticed; another crontab line was "fixed" with `> /dev/null 2>&1` after it flooded an operator's mailbox. What does cron do with whatever a scheduled command prints, and how should a scheduled script's output be handled instead?

level: middleimportance: must knowfreq 65%

basics

~20 s

cron mails any output a job produces to the crontab owner, or to MAILTO if set, and mails nothing when the job prints nothing. Appending > /dev/null 2>&1 discards errors too, which is why failures go unnoticed.

open as a page

A colleague asks why the CI pipeline runs ShellCheck over every shell script when the scripts already work in production. What class of bugs does ShellCheck actually find, how does it decide which shell dialect to apply, and what does its severity scale mean?

level: middleimportance: must knowfreq 62%

basics

~20 s

ShellCheck statically analyses shell scripts and flags quoting, expansion and exit-status mistakes that run without any error message yet behave wrongly. It grades findings as error, warning, info or style, and picks the dialect from the script's shebang.

open as a page

Many bash scripts open with `set -euo pipefail`. What does each of the three options change about how the script behaves, and why do teams make it the standard first line?

level: middleimportance: must knowfreq 78%

basics

~20 s

In bash, set -e aborts the script at the first command that returns non-zero, set -u makes expanding an unset variable an error instead of an empty string, and set -o pipefail makes a pipeline report failure from any stage rather than only the last.

open as a page

A bash script does `workdir=$(mktemp -d)` and removes it with `rm -rf "$workdir"` on its last line. On which exits does that cleanup fail to run, and how do you make removal reliable?

level: middleimportance: must knowfreq 58%

basics

~20 s

The last line only runs when the script reaches it. Any early exit, a strict-mode abort, or a signal skips it and leaks the directory. Register the removal once with a cleanup registered on EXIT, immediately after creating the directory.

open as a page

A cleanup script iterates with `for f in $(ls /var/uploads)` and then runs `rm $f`, over a directory where users choose the uploaded filenames. Which filenames break or subvert that loop, and how would you write it correctly?

level: middleimportance: must knowfreq 72%

basics

~20 s

Word splitting and globbing chew the ls output: a name with a space or newline becomes several loop items, one containing an asterisk expands against the directory, and one starting with a dash is read by rm as an option. Loop over a glob and quote.

open as a page

A container's entrypoint is a bash script whose last line is `myserver --config /etc/app.yml`. Stopping the container always takes the full grace period before the process dies, and the server never runs its shutdown routine. What is wrong with that last line, and what is the fix?

level: seniorimportance: must knowfreq 55%

basics

~20 s

The shell stays as process 1 and the server runs as its child, so the runtime's SIGTERM goes to the shell and never reaches the server. Ending the line with exec replaces the shell, making the server process 1.

open as a page

A bash script begins its option loop with `while getopts "hf:v" opt; do`. Which of h, f and v takes a value, where does bash put that value, and how are `-f out.txt`, `-fout.txt` and `-vf out.txt` each parsed?

level: juniorimportance: should knowfreq 52%

basics

~20 s

A trailing colon means that option takes a value: -f does, h and v do not. getopts puts the option letter in opt and -f's value in OPTARG, accepting -f out.txt, -fout.txt and -vf out.txt alike.

open as a page

A backup script that works when you run it from its own directory fails from a crontab entry with "config/settings.conf: No such file or directory". Which working directory does cron give a scheduled command, which one does systemd give a system service, and how should the script locate its own files?

level: juniorimportance: should knowfreq 50%

basics

~20 s

cron starts the command in the user's home directory and systemd starts a system service in the root directory, so relative paths inside the script resolve somewhere unexpected. Use absolute paths, or derive the script's own directory before opening anything.

open as a page

A bash script is run by a CI job that treats a non-zero exit as a build failure. If the script never calls `exit`, what determines the status it returns, and which conventions should you follow when choosing codes?

level: juniorimportance: should knowfreq 50%

basics

~20 s

A bash script that ends without calling exit returns the exit status of the last command it ran, so a trailing echo or log line reports success even after an earlier failure. Zero means success, any non-zero value means failure, and the status is truncated to the range 0-255.

open as a page

In a bash script, what does `tmpfile=$(mktemp)` give you that `tmpfile=/tmp/myscript.$$` does not?

level: juniorimportance: should knowfreq 55%

basics

~20 s

mktemp creates the file itself, atomically, under a random name with owner-only permissions, and prints the path. A $$-based name is merely a predictable string that nothing has created yet, so another user can pre-create or symlink it first.

open as a page

A colleague's bash trace from `set -x` is a wall of lines that all start with `+ ` and you cannot tell which file or line any of them came from. How do you make bash's trace show the source file, line number and function — and why must that PS4 assignment be single-quoted?

level: middleimportance: should knowfreq 44%

basics

~20 s

Set bash's PS4 variable, which prefixes every xtrace line: PS4='+ ${BASH_SOURCE[0]}:${LINENO}:${FUNCNAME[0]:-main}: '. Single quotes are required so the expansions are stored literally and re-evaluated on each trace line; double quotes would expand them once, freezing one line number.

open as a page

Before shipping a bash script you run `bash -n deploy.sh` and it prints nothing. What has bash actually verified, what has it explicitly not done, and which classes of bug does that check still miss?

level: middleimportance: should knowfreq 38%

basics

~20 s

bash -n reads and parses the whole script without executing any of it, reporting only syntax errors such as an unclosed quote or a missing fi. It cannot see runtime problems: missing commands, unset variables, bad flags, wrong logic, or code built at runtime and passed to eval.

open as a page

A bash script that uses `declare -A` and `mapfile` runs fine on your Linux CI runner, but on a colleague's Mac it dies with `declare: -A: invalid option`. Why does the same script behave differently there, and what are your options?

level: middleimportance: should knowfreq 50%

basics

~20 s

macOS still ships bash 3.2.57 as /bin/bash, and declare -A, mapfile and ${var^^} are all bash 4.0 features. Either guarantee a newer bash on the Mac, or rewrite those constructs so they run on 3.2.

open as a page

A shell script prints messages with `echo "$msg"`. Why do portability-conscious authors replace that with `printf '%s\n' "$msg"`, and what actually breaks when the same script runs under a different shell?

level: middleimportance: should knowfreq 42%

basics

~20 s

echo's handling of backslashes and a leading -n differs by shell: dash's echo expands backslash escapes, bash's does not, and a value of -n disappears entirely. printf's format semantics are specified, so printf '%s\n' behaves the same everywhere.

open as a page

ShellCheck flags a line in your bash script that you have decided is genuinely fine. How do you suppress that one finding without turning the check off repository-wide, what scope does the suppression comment have, and when is suppressing one legitimate?

level: middleimportance: should knowfreq 48%

basics

~20 s

Put a # shellcheck disable=SC2086 comment on its own line immediately above the offending command, with a comment saying why. Placed before the first command it silences the whole file instead; repo-wide suppression belongs in a .shellcheckrc.

open as a page

Even with `set -euo pipefail` at the top, experienced bash authors still write `command || die "message"`. What does an explicit die helper add, and what does a good one look like?

level: middleimportance: should knowfreq 45%

basics

~20 s

An explicit die helper turns a silent abort into a diagnostic: it prints a named, actionable message to stderr and exits with a chosen status. It also covers cases set -e cannot see, such as a missing tool, a missing variable, or a command that reports failure while exiting zero.

open as a page

A script refreshes a config file with `curl -sS https://example.com/config.json > /etc/app/config.json`, and readers occasionally load an empty or half-written file. What is happening, and what is the standard shell fix?

level: middleimportance: should knowfreq 38%

basics

~20 s

The redirection truncates the destination the moment it opens it, long before any bytes arrive, so readers see a growing file — or nothing at all if the download fails. Write to a temp file in the same directory and rename it over the destination.

open as a page

You need to run a command over every matching file in a tree that untrusted users can create files in. Compare `find . -type f | xargs cmd`, `find . -type f -print0 | xargs -0 cmd`, `find . -type f -exec cmd {} \;` and `find . -type f -exec cmd {} +`.

level: middleimportance: should knowfreq 52%

basics

~20 s

Plain find piped to xargs splits names on whitespace and gives quotes and backslashes special meaning, so hostile filenames break it. find -print0 with xargs -0 delimits on NUL, which no filename can contain; find -exec cmd {} + gets the same safety without a pipe, and -exec cmd {} ; is safe but runs one process per file.

open as a page

A bash script in a CI job runs with `set -x`, and the trace both corrupts the stderr output another tool parses and prints an API token into a shared build log. How do you route bash's trace somewhere other than stderr, and what does that still leave you responsible for?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Open a file descriptor for the trace and point BASH_XTRACEFD at it — exec 9>trace.log; BASH_XTRACEFD=9; set -x — and bash writes xtrace there instead of stderr, from bash 4.1 onward. The trace still contains expanded secrets, so its destination must be protected and the switch left off by default.

open as a page

A colleague reports that `./deploy.sh production --dry-run` ignores the flag completely, while `./deploy.sh -v production` works. The script parses its options with `while getopts "v" opt`. Explain both behaviours, and what your choices are for supporting `--dry-run`.

level: seniorimportance: should knowfreq 45%

basics

~20 s

getopts stops at the first argument that is not an option, so parsing ends at production and --dry-run is never examined. getopts also handles only single-character options; long options need a hand-written case loop over the arguments, or the Linux-only enhanced getopt(1).

open as a page

A bash deploy script that parses flags with getopts kept running with its default settings when a user passed a mistyped flag, and separately a CI step that runs `./deploy.sh -h` is marked as a failed build. What is wrong with the script's option-parsing contract, and how should it signal a usage error versus a help request?

level: seniorimportance: should knowfreq 36%

basics

~20 s

An invalid option does not end a getopts loop — getopts still returns success — so a script with no error branch runs on defaults. Usage errors belong on stderr with a non-zero status such as 2; a -h help request belongs on stdout with exit 0.

open as a page

A bash script that works on your Linux build server misbehaves on macOS: `sed -i 's/a/b/' notes.txt` fails with a cryptic message about an undefined label, and `date -d '2 days ago'` is rejected outright. Why do identically named commands behave differently, and how would you make the script work on both?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Linux ships GNU versions of sed, date and stat while macOS ships BSD ones. Only the POSIX-specified core is common; extensions like sed -i and date -d differ or do not exist. Stick to the portable subset, or detect which tool you have.

open as a page

A maintenance script that will run as a systemd service currently opens /var/log/maintenance.log itself, writes its own timestamps and truncates the file when it grows too large. What should a script run under systemd do with its output instead, and how does it mark one line as an error?

level: seniorimportance: should knowfreq 38%

basics

~20 s

A script run under systemd should just print to stdout and stderr; systemd captures both into the journal and adds the timestamp, unit and PID. Both streams get the same default priority, so mark errors with a <3> prefix on the line.

open as a page

A bash script does `source "$LIB_DIR/common.sh"` and ShellCheck reports SC1090, "can't follow non-constant source"; a second script that sources `lib/common.sh` by a literal path gets SC1091 instead, followed by a pile of "referenced but not assigned" warnings. What is the difference between those two codes, and how do you get ShellCheck to actually read the library?

level: seniorimportance: should knowfreq 34%

basics

~20 s

SC1090 means the sourced path is built from a variable, so ShellCheck cannot know which file to read; SC1091 means the path is literal but the file was not made available. Fix both with a # shellcheck source= directive and run with -x.

open as a page

A bash script starts with `set -e`, yet a command that fails inside it does not stop the script. In which contexts does bash deliberately ignore `set -e`, and how do you make the failure surface anyway?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Bash suspends set -e wherever a non-zero status is being used as data: the condition of if, while or until, every command in a && or || list except the last, and any command negated with !. A function called from such a position runs its whole body with set -e suspended.

open as a page

A cron script prevents overlapping runs with `if [ -f /var/tmp/job.pid ]; then exit 0; fi` and then writes its own PID into that file. After one run was killed with `kill -9`, the job never ran again. Why did that guard jam, and what mechanism would not have?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Nothing removes a PID file when the process dies abnormally, so the stale file blocks every later run forever. A kernel advisory lock taken with flock cannot go stale: the kernel drops it when the holder's file descriptors close, however the process died.

open as a page

You inherit a repository with several hundred shell scripts and a currently green CI pipeline, and you want ShellCheck enforced going forward. ShellCheck has no built-in baseline feature. How would you roll it out so the build stays green and the finding count can only go down?

level: principalimportance: should knowfreq 28%

basics

~20 s

Enforce it on a checked-in list of already-clean files rather than on the whole tree, add files to that list as they are fixed, and require every new script to join it. Ratchet the severity threshold down over time.

open as a page

showing 1–30 of 38