You own a bash deployment script that colleagues rerun by hand after it fails halfway. What does making that script idempotent actually require, and how do you decide how far to take it?
answer
- same end state, not a clean exit
- which steps accumulate
- regenerate instead of editing in place
- markers record attempts, not state
- rerun and double-run arrive together
basics
~20 sIdempotence is about the end state, not the exit status: a second run must leave the system in the same shape as the first, whether or not the first finished. That means classifying every step and converting the ones that accumulate rather than converge.
solid answer
~50 sStart by defining the target: after the script runs, these files have this content, this user exists, this version is live. Then walk every step and sort it. Some are naturally convergent — `mkdir -p`, `chmod`, `rsync`, an HTTP `PUT`. Some accumulate — appending a line, incrementing a counter, `useradd`, a `POST` that creates a record. The accumulating ones need converting: guard the append with `grep -qxF`, guard the create with an existence check, or replace an in-place edit with a full regeneration published by rename so a half-written file never exists. Where a step reaches a remote system, the real guarantee belongs there — an idempotency key or a conditional request — not in your `if`. Then decide how far to go: reruns and concurrent runs arrive together, so add a lock; and when the branch count outgrows what you can reason about, that is the signal to move to a tool built for convergence.
code
bash · 15 lines#!/usr/bin/env bash
set -euo pipefail
file=./profile.local
line='export PATH=/opt/app/bin:$PATH'
mkdir -p "$(dirname "$file")"
touch "$file"
if grep -qxF -- "$line" "$file"; then
printf 'already present\n'
else
printf '%s\n' "$line" >> "$file"
printf 'added\n'
figo deeper
Know the definition: running the script twice should leave the system exactly as running it once. Recognise the everyday convergent commands such as mkdir -p, and know that appending with >> is the classic step that is not.
Explain how to convert a step: guard an append with grep -qxF, check before creating a user, and prefer regenerating a whole file and publishing it by rename over editing it in place. Be ready to name which of your steps fail on a second run.
Show the operational reasoning: which half-applied states actually hurt, why a rerun usually overlaps the first run so a lock is part of the answer, and why a step that reaches a remote system needs an idempotency key rather than a shell-side existence check.
Own the proportionality call — how much convergence machinery a script earns given the cost of a half-applied state — and the exit criterion: when the guards outgrow the actions, move to a configuration-management tool or a job runner instead of writing one in bash.
## Idempotent means the end state, not a clean exit The test is not "does the second run avoid errors". It is: after running the script once, then again, is the observable state of the system identical to running it once? A script that fails loudly on the second run is *not* idempotent, but neither is one that exits 0 while quietly appending the same line to a config file for the fifth time. State the target as a description of the world — this directory exists with this owner, this file has these contents, this symlink points at this release — and idempotence becomes a checkable property rather than a vibe. ## Classify every step Walk the script and put each operation into one of three buckets. **Naturally convergent.** `mkdir -p`, `chmod`, `chown`, `rsync --delete`, writing a whole file, an HTTP `PUT` to a known URL. These describe a desired state and reaching it twice costs nothing. **Accumulating.** `>>` appends, `sed -i` edits that match again after their own output, counters, `git clone` into an existing directory, `tar` extraction over live files, an HTTP `POST` that creates a new record each time. Run twice, you get two of something. **Conditionally fatal.** `useradd`, `ln -s`, `mv` of a source that only existed the first time, `createdb`. These fail on the second run, which at least is loud — but under `set -e` a loud failure halfway through is exactly the half-applied state you are trying to escape. ## Converting the awkward ones The usual moves, in rough order of preference: **Regenerate rather than edit.** The single most effective change is to stop editing files in place. Render the whole file from a template into a staging file, validate it, and publish it with a rename. That is idempotent by construction, safe to interrupt, and a diff against the live file tells you whether anything will change at all. **Guard the append.** ```bash line='export PATH=/opt/app/bin:$PATH' grep -qxF -- "$line" "$file" || printf '%s\n' "$line" >> "$file" ``` `-q` is quiet, `-x` matches whole lines, `-F` treats the pattern as a fixed string, and `--` stops a leading dash being read as an option. Note the temptation to "fix" it with `>` instead — that truncates the whole file and trades a duplicate line for data loss. **Check before creating.** `id -u appuser >/dev/null 2>&1 || useradd appuser`. Accept that check-then-act is racy against a concurrent run and cover that with a lock rather than pretending the check is atomic. **Push the guarantee to the system that owns the state.** If a step calls a remote API, an idempotency key or a conditional request (`If-Match`, an upsert, a unique constraint) gives a real guarantee that no amount of shell branching can. Your script should be *carrying* the key, not simulating the semantics. ## Markers and locks are a last resort A sentinel file — `/var/lib/app/.step3-done` — is tempting for a step you genuinely cannot make convergent. It is weak, and it is worth being able to say why: it records that the step was *attempted*, not that the system is still in the resulting state. Someone deletes the user your marker says you created, and the script now skips the repair forever. It is the stale-lock problem in a different costume. Use markers for genuinely one-way operations, put the real state check first where one exists, and make them easy to clear. ## Rerun and concurrent run arrive together The practical reason to care is that a human reruns the script *because* it looked stuck — often while the first copy is still working. So the same effort that makes a rerun safe should also make a concurrent run impossible: take a lock at the top with `flock`, and decide deliberately whether the second copy waits or exits. Reporting matters here too. A rerun should print what it changed and what it found already correct, so an operator can tell "nothing to do" from "did nothing". ## Knowing when bash is the wrong tool This is the part that separates a principal-level answer. Every convergence guard is a branch, and branches multiply: a script with thirty steps and thirty guards has a state space nobody reviews honestly, and the guards themselves become the bug. When you notice you are writing a small configuration-management engine — resource declarations, dry-run mode, change reporting — stop and use one, or move the work to a job runner that provides retries and deduplication as a platform property. Equally, resist the opposite failure: a five-line script that copies one file does not need a marker directory and a state machine. The judgment is proportionality — how expensive is a half-applied state, how often does the script actually fail partway, and who is holding the terminal when it does.
- Why is a sentinel file such as .step3-done a weak idempotency mechanism?It records that the step ran, not that its effect still holds. If someone removes the user or directory the marker vouches for, every later run skips the repair and the drift is permanent. It also has to be cleaned up and versioned like any other state. Prefer checking the real world; keep markers for genuinely one-way operations and make them trivial to clear.
- A step calls a remote API that creates a record. How do you make that safe to retry?Not in the shell. Send an idempotency key — a stable identifier derived from the request, not a fresh random one per run — so the server recognises a repeat and returns the original result, or use an upsert with a unique constraint. A shell-side "check whether it exists first" is a race, not a guarantee; the system that owns the state is the only place the guarantee can live.
- How does making a script rerunnable interact with concurrency?They are the same requirement in practice, because people rerun a job while the first copy is still running. Every check-then-act guard you add is racy against that second copy, so pair idempotence with a lock — `flock` at the top of the script — and decide whether a second run waits or exits. Idempotence without exclusion just means both copies corrupt the state politely.
- What tells you a script has outgrown this approach?When the guards outnumber the actions, when you find yourself building a dry-run mode and change reporting, or when the ordering between steps needs its own diagram. At that point you are writing a configuration-management engine badly; use one, or move to a job runner that provides retries and deduplication. The counter-signal matters too — a short script does not need a state machine.
saying these in an interview costs you the question
- Says idempotent means the second run exits 0
- Replaces an append with > to avoid duplicates
- Trusts a marker file over checking real state
- Ignores that a rerun often overlaps the first run
- Adds guards to every step without questioning the tool