skip to content

A bash deploy script fans work out to background jobs on a shared CI runner. How do you decide the concurrency limit, and would you abort the remaining jobs at the first failure or let everything run and aggregate the failures?

level: principalimportance: nice to knowfreq 24%

answer

  1. two decisions, not one
  2. the bottleneck is usually the far end
  3. the core count a container reports is a lie
  4. blast radius decides the policy
  5. killing the rest is part of failing fast

basics

~20 s

Set the limit from the bottleneck resource — cores for CPU work, the remote rate limit or connection pool for network work, memory per child for everything — not from a core count the container may not actually own. Choose fail-fast for mutating work and aggregate reporting for independent checks.

solid answer

~50 s

Two decisions, and they are independent. The **limit** comes from whichever resource saturates first: cores for compute, but for a deploy it is usually the far end — an API rate limit, a database connection pool, a registry — and memory per child on a container that may have a fraction of the host. `nproc` is a poor default inside a container, since it reflects the CPU affinity mask, not a CFS quota, so on a 64-core host with a 2-core quota it still reports 64. The **failure policy** follows from what the jobs touch: for independent read-only work such as lint or verification fan-out, run everything and print one complete report, because a second run to find the second failure is pure latency. For mutating work, fail fast — stop feeding new jobs, kill the in-flight PIDs, wait for them, and be explicit in the log about what completed, because partial deploys are the real hazard.

go deeper

for a junior

Know that a fan-out needs a limit at all, and that the script's exit status should be non-zero if any job failed. You are not expected to choose the number yet.

for a middle

Be able to reason about which resource saturates first — cores, memory, or the remote service — and to explain why continuing after a failure is fine for checks but risky for anything that writes.

for a senior

Show that a container's reported core count can be wrong, size from measurement rather than habit, and implement fail-fast properly by killing and reaping the in-flight jobs before exiting.

for a principal

Own the tradeoff end to end: blast radius versus feedback latency, what partial completion means to whoever is on call, whether the jobs are idempotent enough for any policy to be recoverable, and when the fan-out should leave bash for the CI matrix or a queue.

## These are two separate questions Candidates usually collapse "how many" and "what on failure" into one answer. They are driven by different things: the limit is a resource question, the policy is a blast-radius question. ## Choosing the limit Start from the bottleneck, not from a habit. - **CPU-bound work** (compiling, compressing, image processing): the ceiling is roughly the number of cores you actually get. More children than cores adds context-switching and memory pressure without adding throughput. - **IO- or network-bound work** (uploads, API calls, remote deploys): the local machine is rarely the constraint. The real cap is on the far end — a rate limit, a connection pool, a per-account concurrency ceiling — and exceeding it converts your parallelism into 429s and retries, which is slower than running fewer jobs. Ask what the far end tolerates before asking what the runner tolerates. - **Memory** is the constraint people skip. N children at 400 MB each is 400 MB × N against a container limit that may be 2 GB. The failure mode is a child returning 137 from the OOM killer, which looks like a flaky job rather than a sizing mistake. - **Shared runner citizenship**: on a machine that runs other people's jobs, your limit is not "what I can get away with" but "what leaves the host usable". A fan-out that saturates a shared runner turns your green build into everyone else's flake. The container trap deserves naming explicitly. `nproc` reports the number of CPUs the process is *allowed to run on*, derived from its affinity mask. A CFS quota (`--cpus=2`, or a Kubernetes CPU limit) does not change the affinity mask, so `nproc` on a 64-core host still says 64 while the cgroup throttles you to two cores' worth of time. Deriving `-P "$(nproc)"` from that produces 64 children fighting over 2 cores. If the limit matters, read it from the environment your platform provides or pin it explicitly, and make it a variable the operator can override rather than a constant buried in the script. ## Choosing the failure policy **Aggregate (run everything, report all failures)** fits work that is independent and side-effect-free: linting 40 packages, verifying 200 config files, checking 50 endpoints. The reason is human latency — if three checks are broken, finding them one build at a time costs three round trips. The cost is runner time spent on jobs whose result you already know will not save the build. **Fail fast (stop at the first failure)** fits work that mutates something. If job 3 of 20 fails halfway through a rolling deploy, continuing means 17 more regions changed while one is known-broken. Fail fast limits the blast radius and gives you a smaller, clearer state to reason about. Fail fast is more work to implement honestly, and half-done versions are worse than none. The full shape is: stop feeding new items, `kill` the PIDs still running, `wait` for each of them so nothing is left orphaned, and then report exactly which items completed, which were aborted, and which never started. A script that exits on the first failure while leaving eight children running has produced the worst of both — a red build *and* uncontrolled mutation still in flight. The honest middle position for deploys is fail-fast plus **idempotency**: every job must be safe to re-run, so recovery is "fix and run again" rather than "work out by hand what got half-applied". If the jobs are not idempotent, no failure policy saves you; that is a design defect one level up. ## What the caller sees Whichever policy you pick, the script's own exit status is one bit and the log carries the detail. Decide deliberately: - non-zero if *any* job failed is the safe default; - the log must name the failing items, not just the count, because the next reader is someone at 2am; - partial success needs to be stated as partial success — "12 of 20 applied, 1 failed, 7 not attempted" — rather than implied by a bare non-zero exit. ## When to stop doing this in bash A `wait -n` pool is genuinely fine for tens of independent, uniform, local jobs. Push past that and the requirements start arriving that bash has no good answer for: per-job retry with backoff, resuming a partially completed run, cross-host scheduling, structured per-job results, timeouts per job rather than per script. Each of those is another hundred lines of fragile shell. When two or three of them are on the table, the fan-out belongs in the CI system's own matrix, a job queue, or a small program in a language with a real concurrency library — and the bash script's remaining job is to invoke it.

  • Why is `-P "$(nproc)"` a poor default inside a container?
    `nproc` reports the CPUs the process is allowed to run on, which comes from its affinity mask. A CFS quota such as `--cpus=2` throttles CPU *time* without touching that mask, so on a 64-core host nproc still says 64 and you fork 64 children into two cores' worth of budget. Read the platform's declared limit, or make the number an overridable variable.
  • What does a correct fail-fast implementation have to do beyond exiting?
    Stop feeding new items, kill the PIDs still in flight, wait for each of them so nothing is orphaned, and report which items completed, which were aborted and which never started. Exiting while eight children keep mutating things is worse than not failing fast at all.
  • When would you argue against parallelising the fan-out in bash at all?
    When per-job retries with backoff, resuming a partial run, per-job timeouts, or cross-host scheduling are on the table. Each is hard to do correctly in shell, and together they mean the fan-out belongs in the CI system's job matrix or a small program in a language with real concurrency support.

saying these in an interview costs you the question

  • Picks the limit from nproc without checking the container quota
  • Treats fail-fast as always the responsible choice
  • Exits on first failure leaving children still running
  • Ignores per-child memory when sizing concurrency
  • Assumes the runner is theirs alone to saturate

context