skip to content

A nightly rsync job over SSH copies the entire dataset every run, although almost nothing changes on the source. What rule does rsync use to decide a file needs transferring, and what are the usual reasons it keeps deciding yes?

level: seniorimportance: should knowfreq 45%

answer

  1. two attributes decide, not the contents
  2. the letter that vanished from the flags
  3. filesystems that round the clock
  4. -ni prints which attribute differs
  5. the fix that reads every byte on both sides

basics

~20 s

By default rsync skips a file only when its size and modification time match on both sides. Anything that stops mtime matching — not passing -t or -a, a destination filesystem with coarse timestamp resolution, or a source process rewriting files — makes every file look changed. Run with -ni to see which attribute differs.

solid answer

~50 s

rsync's default file selection is the "quick check": transfer the file unless its size **and** modification time both already match on the destination. So a job that re-copies everything is telling you that mtime is not matching, and there are three usual causes. First, the flags — `rsync -r` or `-rlpgoD` without `-t` does not preserve times, so every file lands with a fresh mtime and looks changed on the next run; `-a` includes `-t`, which is most of why `-a` exists. Second, the destination cannot represent the source's timestamps: FAT/exFAT and some SMB mounts round to two seconds, which is what `--modify-window=1` (or 2) is for. Third, the source is genuinely being rewritten — a build or an export regenerating identical content with new mtimes, in which case rsync is right and `-c` (`--checksum`) or fixing the producer is the answer. Diagnose it with `rsync -ni`: the itemised change string names the attribute that differs.

go deeper

for a junior

Know that rsync decides what to send by comparing size and modification time, not by reading the files, and that -a is the option that preserves times so the next run can skip them.

for a middle

Explain the quick check precisely and connect it to the flags: no -t means every file lands with a fresh mtime and looks changed. Name --checksum, --size-only and --ignore-times and say what each substitutes for the default rule.

for a senior

Lead with diagnosis. Read the -ni itemise string to identify the differing attribute, stat both sides before theorising, and weigh --checksum's full-read cost against the redundant transfer instead of reaching for it reflexively.

for a principal

Treat it as a data-movement design question: whether the producer should stop rewriting unchanged output, whether the sync should read from an immutable artefact instead, and what the nightly window and link budget can actually support as the dataset grows.

## The rule rsync's default selection test is deliberately cheap and is called the *quick check*: a file on the destination is considered up to date if its **size** and its **modification time** both match the source. If either differs, the file is queued for transfer. Nothing is read, no checksum is computed — this is a metadata comparison, which is why rsync can walk millions of files quickly. Everything else follows from that one sentence. A job that re-transfers unchanged data is not a bug in the algorithm; it is a report that mtimes are not matching. ## Diagnose before you theorise Run the job with `-n -i` (`--dry-run --itemize-changes`) and read the change strings. The format is a per-attribute string such as `>f.st......`: - position 1: the action (`>` received, `<` sent, `c` created, `.` no change) - position 2: the file type (`f` file, `d` directory, `L` symlink) - then one letter per attribute that differs: `c` checksum/content, `s` size, `t` time, `p` permissions, `o` owner, `g` group `>f..t......` means the content is fine and only the timestamp differs — that is the fingerprint of this whole failure mode. `>f+++++++++` means the destination has no such file at all, which points somewhere else entirely (wrong path, wrong slash, destination being wiped between runs). Reading this string first saves an hour of guessing. ## Cause 1: the flags do not preserve times `-r` alone gives you recursion and nothing else. Times are preserved by `-t`, which is bundled into `-a`. Without it every file lands with the time it was written at the destination, which is by definition *now*, so on the next run all mtimes differ and the whole tree transfers again. The same happens if someone "tidied" a command into `-rlpgoD` (archive minus `-t`) or added `--no-t` to `-a`. This is the single most common answer, and it is worth stating the causality out loud: `-a` is not just about fidelity, it is what makes repeat runs cheap. ## Cause 2: the destination cannot hold the timestamp Some filesystems store coarser times than Linux does. FAT and exFAT round to two seconds; some SMB/CIFS mounts and some object-store gateways behave similarly, and a few network filesystems apply their own clock. rsync then writes the right value and reads back something slightly different, so the comparison fails forever. The fix is `--modify-window=NUM`, which tells rsync to treat times within NUM seconds of each other as equal — `--modify-window=1` is enough for FAT's two-second granularity, and `--modify-window=2` is a common belt-and-braces setting. A related variant: a destination mounted with an option that ignores utimes, or a container/NAS layer that stamps its own times. Same symptom, same fix, but confirm with `stat` on both sides before reaching for the flag — you want to see the actual pair of timestamps, not assume. ## Cause 3: the source really is changing A build that regenerates identical bytes, an export that rewrites a whole directory nightly, or a `cp` step upstream that resets times all produce files whose content is unchanged but whose mtime is new. rsync is behaving correctly; the producer is lying about change. Two options: - **`-c` (`--checksum`)** replaces the quick check with a full-file checksum comparison on both sides. It is correct and it is expensive: both hosts read every byte of every file before anything is transferred. On a large dataset over a nightly window this can cost far more than the redundant transfer it avoids. Use it deliberately. - **Fix the producer** so unchanged outputs keep their timestamps, or sync from a source of truth that does not rewrite. This is usually the better engineering answer. ## The inverse trap The mirror image of this failure is a job that transfers *too little*. `--size-only` compares size alone and ignores mtime entirely; it is sometimes reached for as a "fix" for re-transfer problems, and it silently misses every same-size edit — a flipped byte in a fixed-length record, an image edited in place, a config value swapped for another of equal length. If you find `--size-only` in an existing job, treat it as a defect to justify, not a setting to keep. `-I`/`--ignore-times` is the opposite extreme: it disables the quick check and transfers everything, which is occasionally useful for a one-off forced refresh and disastrous as a permanent setting. ## Checklist for the incident 1. `rsync -ni ...` and read the itemise string — is it `t`, `s`, `c`, or `+++++++++`? 2. If `t`: is `-t`/`-a` present? `stat` the same file on both sides and compare timestamps. 3. If the timestamps differ by a sub-second or two-second amount: `--modify-window`. 4. If content genuinely changed: look upstream at what rewrites the files. 5. Only then consider `-c`, and measure what the full read costs on both hosts.

  • Why not just add `--checksum` and be done with it?
    Because `-c` makes both hosts read every byte of every file before deciding anything, converting a metadata walk into a full-dataset read on two machines. On a large tree that can exceed the nightly window and cost more than the redundant transfer. It is the right answer when the source genuinely rewrites unchanged content and you cannot fix the producer — measure the read cost first.
  • Someone has set `--size-only` on the job. What are you worried about?
    Same-size edits are invisible. A flipped byte in a fixed-width record, an in-place image edit, or a config value swapped for another of the same length will never transfer, and the mirror silently diverges. `--size-only` is defensible only where mtimes are known-meaningless and content is known-append-only. Otherwise it trades a performance annoyance for a correctness bug.
  • How would you confirm a timestamp-granularity problem rather than guessing?
    Pick one file rsync insists on re-transferring and run `stat` on it on both hosts, comparing the modification times directly. A difference of a second or two, or a destination time that is always rounded, points at filesystem granularity and `--modify-window`. A destination time equal to the last run's clock time points instead at times not being preserved at all.

saying these in an interview costs you the question

  • Assumes rsync compares file contents by default
  • Blames the network for a re-transfer of unchanged data
  • Sets --size-only to stop repeated transfers
  • Adds --checksum without weighing the full read cost
  • Thinks -r and -a differ only in metadata fidelity

context