You are hardening a nightly `rsync -a --delete` mirror to a remote host over SSH: it is regularly killed part-way through, it saturates the uplink, and once it emptied the destination. Which rsync options make the job resumable, bounded and safe to interrupt?
answer
- four failures, four flags
- where the half-written file should live
- a ceiling on how much it may remove
- prune at the end, not at the start
- an empty source is not an error
basics
~20 sUse --partial-dir so a killed run resumes instead of restarting large files, --bwlimit to cap throughput, --max-delete as a circuit breaker, --delete-after or --delay-updates so the destination is not pruned before the new data lands, and --timeout to fail a stalled run. Verify the source is really mounted before running.
solid answer
~50 sFour separate problems, four separate answers. **Resumability**: `--partial-dir=.rsync-partial` keeps partly transferred files in a hidden directory so the next run continues them instead of starting the largest file over, and it keeps the incomplete data out of the live tree (better than plain `--partial`); `-P` is the interactive shorthand for `--partial --progress`. **Bandwidth**: `--bwlimit=20M` caps throughput, and `--timeout=600` makes a stalled transfer die instead of hanging until morning. **Deletion safety**: `--max-delete=1000` aborts with a non-zero status if the run wants to remove an implausible number of files, and `--delete-after` (or `--delay-updates`) means the pruning happens once the new content is in, so an interrupted run leaves the destination populated rather than gutted. **The wipe**: rsync aborts on an unreadable source, but an *empty* source that mounts cleanly is not an error — so the wrapper must check a sentinel file before rsync ever runs. And keep a `--link-dest` snapshot chain so even a correct-looking bad run is recoverable.
code
bash · 10 lines#!/bin/sh
set -e
# rsync cannot tell an unmounted source from an empty one -- check first
test -f /srv/data/.mirror-sentinel || { echo "source not mounted" >&2; exit 1; }
rsync -a --delete --delete-after --max-delete=1000 \
--partial-dir=.rsync-partial --bwlimit=20M --timeout=600 \
--link-dest=/srv/mirror/previous --stats \
-e 'ssh -p 2222' \
/srv/data/ backup@host:/srv/mirror/current/go deeper
Know that -P resumes and shows progress, and that --delete removes destination files with no source counterpart — so a mirror command deserves a --dry-run before it is trusted.
Match each option to the failure it addresses: --partial-dir for resumability, --bwlimit and --timeout for the link, --max-delete for deletion blast radius, and --delete-after for the state an interrupted run leaves behind.
Show the operator's instincts: the empty-but-mounted source that rsync cannot detect, checking the exit status so the guards actually fire, --link-dest snapshots so a bad run is recoverable, and knowing what --delete-after costs on a very large tree.
Argue the recovery objective first and let it choose the mechanism. Decide when file-level synchronisation is the right primitive at all versus filesystem or block-level snapshots and object versioning, who may run destructive jobs unattended, and how restores are tested rather than assumed.
## Separate the failure modes The job has three distinct defects and they need three distinct fixes. Answering "add `-P`" to all of them is the weak answer. ## 1. It restarts from scratch after being killed By default rsync writes each file to a temporary in the destination directory and renames it on completion; if the run dies, that temporary is deleted and the next run starts the file over. For a mirror containing multi-gigabyte files on a link that cannot finish one in a night, the job never converges. - **`--partial`** keeps the partial file so the next run can use it as a starting point. - **`--partial-dir=DIR`** is the better production form: partial data goes into a named directory (relative paths are per-destination-directory, e.g. `--partial-dir=.rsync-partial`) instead of sitting in the live tree under the real filename. Nothing downstream ever sees a truncated file wearing the correct name. It also implies `--partial`. - **`-P`** equals `--partial --progress` and is for humans at a terminal; in a scheduled job use `--partial-dir` and `--info=progress2` if you want a single summary progress line in the log. - `--append-verify` is *not* a general resume option: it assumes the destination is a prefix of the source and verifies only that prefix. Right for append-only logs, wrong and quietly destructive for files that are rewritten. ## 2. It saturates the uplink and can hang - **`--bwlimit=RATE`** caps the transfer rate (`--bwlimit=20M`). Cruder than proper traffic shaping, but it is one flag and it works. - **`--timeout=SECONDS`** aborts if no data moves for that long, which converts a wedged run into a failed run you can alert on. **`--contimeout=SECONDS`** bounds the initial connection separately. - `-z` is worth a thought rather than a reflex: on already-compressed data it costs CPU on both ends for nothing. `--skip-compress=` excludes suffixes, `--compress-choice=` picks the algorithm. - Since the transport is SSH, `-e 'ssh -p 2222'` is how you pass a non-default port or other client options through to the connection. ## 3. It once emptied the destination This is the serious one, and rsync's own protections only cover part of it. - **`--max-delete=N`** is the circuit breaker. rsync stops deleting once N removals are reached and exits with a non-zero status, so an anomalous run fails loudly instead of succeeding destructively. Set it from the tree's real churn — a mirror that normally removes a few dozen files a night has no business removing ten thousand. - **Deletion timing.** With `--delete`, rsync 3.x prunes during the transfer by default (`--delete-during`). `--delete-after` moves all removals to the end, so an interrupted run has added the new content but not yet removed the old — a strictly safer intermediate state for a mirror something else is reading. `--delete-before` is the old behaviour and is the worst choice here: it empties first and requires building the whole file list up front rather than incrementally. Note the cost of `--delete-after`: it needs the full list too, so it gives up incremental recursion's earlier start and lower memory on very large trees. - **`--delay-updates`** puts every updated file in a holding area and renames them all at the end, so the destination flips much closer to atomically. It costs temporary space equal to the changed data. - **The unmounted-source trap.** rsync will abort rather than delete everything if it hits an I/O error reading the source (that is what `--ignore-errors` overrides — never set it on a `--delete` job). But a backup source whose filesystem failed to mount is usually an *empty, perfectly readable* directory, which is not an error. rsync then does exactly what it was told and `--delete` empties the mirror. The guard cannot come from rsync: the wrapper must verify a sentinel file exists in the source, or that `mountpoint -q` succeeds, before invoking it. - **Keep history.** `--link-dest=DIR` builds the new tree with hard links to an unchanged previous snapshot, so each night costs only the changed data yet every night is a complete browsable tree. A mirror alone replicates mistakes; a snapshot chain lets you walk back to the night before. ## 4. Make the run auditable - `-n -i` (`--dry-run --itemize-changes`) before any change to the command, read in full. - `--stats` in the log so you can see literal versus matched bytes and the file counts over time. - **Check the exit status.** rsync's exit codes are meaningful — notably 23 and 24 for partial transfers, and 25 for hitting `--max-delete` — and a job whose wrapper ignores the status has no safety at all, since every option above signals failure by exiting non-zero. ## A defensible nightly command ``` rsync -a --delete --delete-after --max-delete=1000 \ --partial-dir=.rsync-partial --bwlimit=20M --timeout=600 \ --link-dest=/srv/mirror/previous --stats \ -e 'ssh -p 2222' \ /srv/data/ backup@host:/srv/mirror/current/ ``` Every flag there answers a specific failure you can name, which is the standard to hold any production rsync command to.
- Why prefer `--partial-dir` over plain `--partial` in a scheduled job?`--partial` leaves the incomplete file in place under its real name, so anything reading the destination can pick up a truncated file that looks finished. `--partial-dir` puts the partial data in a named side directory, keeping the live tree free of half-written files while still letting the next run resume. It implies `--partial`, so you do not need both.
- What does `--delete-after` cost you compared with the default `--delete-during`?It has to build the complete file list before it can know what is extraneous, which gives up rsync 3.x's incremental recursion — so the transfer starts later and uses more memory on a tree with millions of files. You buy a safer interrupted state: new content is already in place and nothing has been removed yet. On huge trees, weigh it.
- rsync aborts rather than deleting everything when the source is unreadable. Why is that not enough protection?Because the realistic failure is not an I/O error. A backup source whose filesystem failed to mount presents as an empty directory that reads perfectly, so rsync sees a legitimate empty source and `--delete` removes the entire mirror. Nothing inside rsync can distinguish that from a real deletion — the wrapper has to check a sentinel file or the mount state before running.
- Why does the wrapper need to inspect rsync's exit status?Because every safety mechanism here reports by exiting non-zero and otherwise stays quiet. `--max-delete` being hit, a timeout, a partial transfer — all leave a job that looks finished if you ignore the status. Codes 23 and 24 indicate partial transfers and 25 indicates the delete limit was reached; a wrapper that always exits 0 has turned the guards off.
saying these in an interview costs you the question
- Answers every failure mode with -P
- Uses --ignore-errors on a job that deletes
- Assumes rsync detects an unmounted source directory
- Treats a --delete mirror as a backup with history
- Ignores rsync's exit status in the scheduled wrapper