skip to content

A bash script must run the same task over 500 inputs but never have more than 8 running at once. How do you bound the concurrency, and how do you still find out which inputs failed?

level: seniorimportance: should knowfreq 42%

answer

  1. never fork all 500 at once
  2. a flag that caps concurrent invocations
  3. backfill as each one finishes
  4. one exit code for the whole batch
  5. the worker records its own failure

basics

~20 s

Two idioms bound it: pipe the inputs into xargs -P 8 -n1, or keep a PID list in bash and call wait -n whenever 8 jobs are in flight. Neither reports which input failed unless the worker records failures itself.

solid answer

~40 s

The quick answer is `printf '%s\n' "${inputs[@]}" | xargs -P 8 -n1 ./task.sh` — xargs keeps at most eight invocations alive and starts a new one as each finishes. Its exit status is 123 if any invocation exited between 1 and 125, which tells you *something* failed but not what, so the worker itself should append the failing input to a failures file. The pure-bash alternative is a loop that backgrounds each item, appends `$!` to an array, and calls `wait -n` (bash 4.3+) whenever the in-flight count reaches 8 — more code, but you keep the PID-to-input mapping and can report per-item results. Give each child its own log file either way, because eight concurrent writers interleave stdout unreadably.

code

bash · 12 lines
bash
#!/usr/bin/env bash
set -uo pipefail

# Cap at 8 concurrent workers; each records its own failure.
printf '%s\n' item-{1..500} \
  | xargs -P 8 -n1 sh -c './task.sh "$1" >"logs/$1.log" 2>&1 || echo "$1" >>failures.txt' _

if [[ -s failures.txt ]]; then
  echo "failed items:" >&2
  cat failures.txt >&2
  exit 1
fi

go deeper

for a junior

Know that backgrounding one job per input does not scale and that xargs -P exists to cap how many run at once. Being able to write the one-line xargs form is enough at this level.

for a middle

Explain how xargs backfills as invocations finish, what -n1 versus -I{} changes, and why the batch exit status cannot identify the failing input. Know that wait -n is the bash-side equivalent building block.

for a senior

Show the full pool loop with per-item failure reporting, name the stdin-sharing and output-interleaving traps, and justify the concurrency number from the actual bottleneck rather than picking 8 by habit.

for a principal

Own the policy: fail-fast versus complete-report, what partial completion means for downstream state, whether an extra dependency like GNU parallel is acceptable in your images, and at what scale this stops being a shell script's job.

## Why the bound matters `for i in "${inputs[@]}"; do task "$i" & done; wait` forks 500 children at once. On a laptop that is merely slow; on a shared CI runner or against a rate-limited API it is an outage. Bounded fan-out means: at most N in flight, start a replacement as each one finishes. There are two standard implementations, and a senior answer names both and says when each fits. ## Option 1 — `xargs -P` ```bash printf '%s\n' "${inputs[@]}" | xargs -P 8 -n1 ./task.sh ``` `-P 8` sets the maximum number of concurrent invocations; `-n1` gives each invocation exactly one argument. (`-I{}` also forces one argument per command and lets you place it anywhere in the template, at the cost of not batching.) On GNU xargs, `-P0` means "as many as possible", which is exactly the thing you were trying to avoid. What xargs gives you: a one-liner, a hard concurrency cap, and automatic backfill. What it costs you: - **Status granularity.** GNU xargs exits **123** if any invocation exited with a status between 1 and 125, 124 if one exited 255, 125 if one was killed by a signal, 126 if the command could not be run, and 127 if it was not found. So you learn the *class* of failure and nothing about which input caused it. - **Identity.** Recover it inside the worker: `xargs -P 8 -n1 sh -c 'task "$1" || echo "$1" >>failures.txt' _` — the trailing `_` fills `$0` so `$1` is the input. - **Shared stdin.** All children inherit the same stdin, which is the pipe xargs is reading. A child that reads stdin (`ssh` is the classic offender) eats the input list. Use `ssh -n`, or feed the list with `xargs -a list.txt` and leave stdin free. - **Argument safety.** With inputs that can contain spaces or newlines, generate them NUL-separated and use `xargs -0`; the general hostile-input handling is its own topic, but the habit belongs here too. ## Option 2 — the bash worker pool ```bash max=8 declare -A name_of failed=() for item in "${inputs[@]}"; do while (( ${#name_of[@]} >= max )); do wait -n -p done_pid || failed+=("${name_of[$done_pid]}") unset 'name_of[$done_pid]' done ./task.sh "$item" & name_of[$!]="$item" done while (( ${#name_of[@]} )); do wait -n -p done_pid || failed+=("${name_of[$done_pid]}") unset 'name_of[$done_pid]' done (( ${#failed[@]} == 0 )) || { printf 'failed: %s\n' "${failed[@]}" >&2; exit 1; } ``` `wait -n` blocks until *some* child finishes and returns its status; `-p done_pid` (bash 5.1+) tells you which one, so you can map it back to the input. Without 5.1, keep the PID array and drop finished entries by polling, or simply record failures inside the worker as in option 1. A pre-4.3 fallback that still works on macOS's bash 3.2 counts jobs instead of using `wait -n`: ```bash while (( $(jobs -r | wc -l) >= max )); do sleep 0.2; done ``` It is a poll loop rather than an event, so it is coarser, but it holds the bound. ## Choosing between them - Reach for **xargs -P** when the work is uniform, per-item identity can live inside the worker, and you want the script to stay small. It is also the option that survives being handed to someone who does not read bash. - Reach for the **bash pool** when you need per-item reporting in the parent, different arguments per job, or a fail-fast policy (on the first non-zero `wait -n`, kill the remaining PIDs and stop feeding new ones). - **GNU parallel** does both plus grouped output — it buffers each job's output and prints it as one block instead of interleaving — but it is an extra dependency that is frequently absent from minimal container images, so check before you depend on it. ## Output and observability Eight concurrent writers to the same stdout interleave at arbitrary boundaries; partial lines from different jobs can end up on the same line. Redirect each child to its own file (`./task.sh "$item" >"logs/$item.log" 2>&1 &`) and concatenate afterwards, or prefix every line with the item so the mixed log is still greppable. ## Picking the number 8 is a placeholder. CPU-bound work tops out near the core count; network- or IO-bound work can run far wider but is usually capped by something on the far end — a connection pool, an API rate limit, a disk. Memory is the constraint people forget: N children each taking 300 MB will OOM a 2 GB container long before the CPU saturates.

  • GNU xargs returns 123 after your parallel run. What can you conclude?
    At least one invoked command exited with a status between 1 and 125 — but not how many, and not which input. That is why the worker should record failures itself, for example `sh -c 'task "$1" || echo "$1" >>failures.txt' _`. Distinct codes exist for the other cases: 125 for a signalled command, 126 for one that could not be run, 127 for not found.
  • Why can running ssh under `xargs -P` silently consume the input list?
    Every child inherits xargs's stdin, which is the pipe carrying the remaining items, and ssh reads stdin to forward it to the remote side. It swallows items that were meant for later invocations. Fix it with `ssh -n`, or read the list from a file with `xargs -a list.txt` so stdin is not the list.
  • How would you make this fan-out fail fast instead of running all 500 items?
    Use the bash pool: on the first non-zero return from `wait -n`, stop feeding new items and `kill` the PIDs still in flight, then wait for them so nothing is left running. It is only safe when the work is independent and abandoning it mid-way leaves no partial state behind.
  • Your worker writes to stdout and the combined log is unreadable. What do you change?
    Give each child its own output file and concatenate after the run, or have the worker prefix every line with its item identifier. Eight concurrent writers can interleave mid-line, so nothing in the shell guarantees whole-line atomicity for you.

saying these in an interview costs you the question

  • Forks one background job per input and calls it parallel
  • Thinks xargs -P names which invocation failed
  • Uses -P0 believing it sets a limit
  • Assumes concurrent children's stdout stays line-separated
  • Sets the limit from the host core count in a container

context