A 2 GB file on a remote host changes by a few kilobytes a day, yet `rsync -a` over a slow link finishes in seconds. Describe the delta-transfer algorithm that makes that possible, and name a common situation where rsync does not use it at all.
answer
- the side with the old copy speaks first
- two checksums, one cheap and one expensive
- the window advances one byte, not one block
- insertions shift matches instead of breaking them
- local to local skips the whole thing
basics
~20 sThe receiver splits its existing copy into blocks and sends the sender a weak rolling checksum plus a strong checksum for each. The sender rolls a window byte by byte through its version, and transmits only unmatched literal data plus references to blocks the receiver already has. Local-to-local copies skip this and send whole files.
solid answer
~50 sThe trick is that the side holding the *old* copy does the describing. The receiver divides its existing file into fixed-size blocks and, for each, sends a cheap rolling checksum and a strong cryptographic one. The sender slides a window of that block size byte by byte through its own version; at every offset it updates the rolling checksum in constant time, and only when that cheap value hits the receiver's table does it compute the strong checksum to confirm a real match. Matches become a short reference; everything between them is sent as literal bytes. The receiver then rebuilds the file from its own blocks plus the literals. That is why a few kilobytes of change in 2 GB costs kilobytes on the wire — the byte-by-byte rolling window means insertions do not spoil the match, they just shift it. rsync deliberately turns this off when both paths are local, where `--whole-file` is the default because reading both files costs more than copying one.
code
bash · 7 lines# Force the delta algorithm on a local copy and show what it bought
rsync -a --no-whole-file --stats /data/db.img /mnt/mirror/db.img
# Typical tail of the output:
# Literal data: 3,180 bytes
# Matched data: 2,147,480,468 bytes
# total size is 2,147,483,648 speedup is 674,712.09go deeper
Be able to say that rsync sends only the parts of a file that differ, rather than the whole file, and that this is why it beats a plain copy on repeat runs. The mechanism can wait.
Walk the algorithm: receiver blocks and checksums its copy, sender rolls a window one byte at a time, weak checksum filters, strong checksum confirms, output is block references plus literals. Say why the byte-by-byte roll matters for insertions.
Bring the tradeoff. Delta transfer costs full reads and CPU on both hosts, so it wins only when the link is the bottleneck — which is exactly why local-to-local defaults to --whole-file. Mention --stats and the speedup line as the way to check the assumption.
Judge when rsync is the right synchronisation primitive at all. For very large single files that change constantly, block-level replication, an append-only log or a database's own replication beats re-deriving deltas from scratch on every run.
## The problem the algorithm solves The naive way to update a remote copy is to send the whole file. The clever-looking alternative — "just send the parts that changed" — is impossible on its own, because the sender does not know what the receiver has. Nor can the receiver ask for the changed parts, because it does not know what the sender has. Neither side may read the other's file, and comparing them directly would mean shipping one of them across the link, which is the thing you are trying to avoid. rsync's answer is to have the receiver send a compact *description* of its copy, and let the sender search for that description inside its own. ## Step by step **1. The receiver blocks its existing file.** It splits its copy into non-overlapping blocks of a fixed size — chosen from the file's size by default, adjustable with `-B`/`--block-size`. **2. For each block it computes two checksums.** A *weak* rolling checksum, cheap to compute and cheap to update, and a *strong* checksum (MD5 in current protocol versions) that makes an accidental collision negligible. The list of pairs goes to the sender. It is small: a few dozen bytes per block, not per byte. **3. The sender rolls a window across its own version.** Starting at offset 0, it takes a window of the block size and computes the weak checksum. Then it advances by **one byte** and updates that checksum in constant time — that is the entire point of a *rolling* checksum: the value for the window at offset n+1 is derived arithmetically from the value at offset n, the byte leaving and the byte entering, with no re-reading of the window. **4. Cheap test, then expensive test.** At each offset the sender looks the weak checksum up in a hash of the receiver's table. Almost always there is no hit and it moves on for one byte's worth of arithmetic. On a hit it computes the strong checksum and compares; only a strong match counts. **5. Output.** A confirmed match emits a reference to the receiver's block index and the window jumps forward a full block. Unmatched bytes accumulate and are sent as literal data. The stream to the receiver is therefore a mix of "copy your block 4,102" and "here are 3,180 literal bytes". **6. Reconstruction.** The receiver assembles a new file from its own blocks and the literals, verifies a whole-file checksum, then renames the temporary into place. Because the sender advances one byte at a time rather than block by block, an insertion in the middle of a file does not destroy the matches after it — the window simply re-aligns. That is what separates rsync from a naive block-by-block comparison, and it is the detail interviewers listen for. ## What it costs Delta transfer trades **network** for **CPU and disk on both sides**. Both hosts read the entire file — the receiver to checksum it, the sender to roll over it — and both spend CPU on checksums. The saving is real only when the link is the bottleneck. That is exactly the tradeoff rsync encodes in its defaults: - **When both source and destination are local paths, rsync uses `--whole-file` (`-W`) by default** and does not run the delta algorithm at all. Reading two files to avoid writing one makes no sense on a local disk. `--no-whole-file` forces the algorithm on if you really want it (a slow network filesystem mounted locally is the usual reason). - **There is no delta on a first copy.** With no file at the destination there are no blocks to reference, so the whole thing is sent. Delta transfer is an *update* optimisation, and people benchmarking rsync against scp on a fresh copy see little difference, correctly. - **Files skipped by the quick check never reach this stage.** rsync's file-selection rule — same size and modification time means skip — runs first. Delta transfer only applies to files rsync has already decided need updating. ## Flags that interact with it - `--inplace` writes updates directly into the destination file instead of a temporary. It saves space and preserves hard links and the file's identity, but it means a crash mid-transfer leaves a corrupt destination, and it restricts which blocks can be reused. - `--append` / `--append-verify` assume the destination is a prefix of the source — right for append-only logs, wrong and destructive for anything else. - `-c`/`--checksum` is a different thing entirely: it changes *file selection* (checksum the whole file instead of trusting size and mtime), not how a selected file is transferred. ## Reading the result At the end of a run rsync prints `total size is N speedup is X`. The speedup is total logical bytes divided by bytes actually sent. On a well-matched update it can be in the hundreds; on a fresh copy it hovers just above 1. `--stats` breaks out literal versus matched data, which is the number to quote when someone asks whether the delta algorithm is earning its keep.
- Why does rsync compute two checksums per block instead of one strong one?Cost. The sender tests every byte offset, so the per-offset test must be arithmetic on a rolling value rather than a hash of the whole window. The weak checksum is a filter that discards nearly every offset for almost nothing; the strong checksum runs only on the rare candidate hit and rules out collisions. One strong checksum per offset would make the search hopelessly expensive.
- Why does rsync use whole-file transfer by default when both paths are local?Delta transfer trades network bytes for reading both files and checksumming them. Between two local paths there is no network to save, so the algorithm would read 2 GB twice to avoid writing 2 GB once — a clear loss. `--no-whole-file` overrides it, which is worth doing when the "local" destination is actually a slow network mount.
- Does the delta algorithm run on files that rsync's quick check skipped?No. File selection happens first: rsync skips any file whose size and modification time already match on both sides, and those files are never opened for transfer. Delta transfer applies only to files already chosen for updating. That ordering is why a mistake in the selection rules — no `-t`, say — costs you far more than any tuning of the delta stage.
saying these in an interview costs you the question
- Says rsync compares the two files directly
- Thinks the sender advances one block at a time
- Claims delta transfer helps on a first copy
- Believes the delta runs on local disk-to-disk copies
- Confuses --checksum with how a file is transferred