On a copy-on-write filesystem such as Btrfs or ZFS, what is a snapshot physically, and why do free space and file layout behave in ways that surprise people once snapshots exist?
answer
- a second reference, not a copy
- cost follows the change rate
- deletion frees nothing while pinned
- rewrites land somewhere new
- same device, so not a backup
basics
~20 sA snapshot is a second reference to the same on-disk blocks at a moment in time, created instantly with no data copied. Space is consumed later, as the live copy diverges, which is why deleting files frees nothing while a snapshot still references them.
solid answer
~50 sCopy-on-write filesystems never overwrite a block in place; a modification is written elsewhere and the tree pointing at it is updated. Taking a snapshot therefore costs almost nothing — it just pins the existing tree so those blocks stay referenced. Cost arrives afterwards: every block the live filesystem changes must now exist twice, once for the snapshot and once for the current state. Two consequences follow. Deleting a large file frees no space while any snapshot still references its extents, so `df` on Btrfs is misleading and you want `btrfs filesystem usage` for the real picture. And because rewrites relocate blocks, a heavily overwritten file fragments; Btrfs offers `chattr +C` or the `nodatacow` mount option to opt individual files out, at the price of losing data checksums for them. A snapshot on the same device is not a backup.
code
bash · 3 linesbtrfs subvolume snapshot -r /srv/data /srv/.snapshots/data-2026-08-20
btrfs subvolume list /srv
btrfs filesystem usage /srvgo deeper
Know that a snapshot on Btrfs or ZFS is taken instantly because no data is copied, that it captures the filesystem at a point in time, and that it lives on the same device as the original.
Explain the mechanism: copy-on-write never overwrites in place, so a snapshot is just a second reference to the same extents, and space is consumed afterwards in proportion to how much the live data changes.
Show the operational consequences you have lived with — deletions that free nothing, df numbers that lie, and databases or VM images fragmenting on copy-on-write — and the mitigations, including what nodatacow gives up.
Own the retention policy and its cost model: how many snapshots, for how long, sized against the volume's change rate, plus which snapshots are replicated off-box so the organisation's recovery story does not rest on a single device.
## Copy-on-write in one paragraph A traditional filesystem such as ext4 or XFS writes a modified block back over the block it came from. A copy-on-write filesystem never does. When a block changes, the new version is written to free space, and the metadata node pointing at it is itself rewritten to free space, and so on up to the root of the tree, which is updated last. The effect is that the on-disk state moves atomically from one consistent tree to the next: either the new root is committed or it is not, and there is no window in which a half-updated structure is visible. That is also why these filesystems have no separate data-journalling mode — the write-elsewhere discipline is the crash-consistency mechanism. ## What a snapshot therefore is Once nothing is overwritten in place, a snapshot is almost free. On Btrfs a subvolume is an independent file tree within the filesystem, and a snapshot is a new subvolume that starts out pointing at exactly the same extents as the source: ``` btrfs subvolume snapshot -r /srv/data /srv/.snapshots/data-2026-08-20 ``` No data is copied. The operation is effectively constant time regardless of whether the subvolume holds one gigabyte or one terabyte, because all it does is create a second reference to an existing tree and mark those extents as shared. `-r` makes the snapshot read-only, which is what you want for anything you intend to keep or send elsewhere; without it the snapshot is writable and diverges independently. ZFS expresses the same idea as `zfs snapshot pool/dataset@name` over a dataset. This is why snapshot-based workflows are so appealing: a pre-upgrade snapshot, a roll-back after a failed deployment, a stable point-in-time tree for a backup to read from while the live system keeps writing. ## Where the space actually goes The cost is deferred, not avoided. From the moment the snapshot exists, every block the live subvolume rewrites must exist twice: the old version pinned by the snapshot, the new one in the live tree. Space consumption therefore tracks the *rate of change*, not the size of the snapshot. The consequence that catches people out is deletion. An engineer responding to a full disk deletes a 200 GB file and sees no space returned. That is correct behaviour: the extents are still referenced by a snapshot, so they cannot be freed. Nothing is reclaimed until every reference is gone, meaning the snapshots holding those extents are deleted too. This is also why plain `df` is untrustworthy on Btrfs — a single number cannot express space that is shared between trees, shared metadata, or space allocated to chunks but not yet used. Use the filesystem's own accounting: ``` btrfs filesystem usage /srv btrfs subvolume list /srv ``` ZFS shows the same information per dataset through the `USEDSNAP` and `USEDDS` columns of `zfs list -o space`. Btrfs quota groups can attribute space per subvolume, at a real performance cost that many operators decline to pay. ## Fragmentation and the overwrite workload The second surprise is layout. Because a rewrite of the middle of a file allocates a new block somewhere else rather than reusing the old one, a file subjected to sustained random overwrite scatters across the device. For append-mostly or write-once data this barely matters. For a database's page files or a virtual-machine disk image it matters a great deal: after months of random writes, a scan that ought to be sequential becomes thousands of seeks, and on rotational media the difference is dramatic. Btrfs gives you an opt-out. `chattr +C` sets the no-copy-on-write attribute on a file, and the `nodatacow` mount option applies it filesystem-wide. Two caveats make this less of a free win than it sounds. The attribute only takes effect on a file with no data yet, so in practice you set it on the *directory* before the files are created and let new files inherit it. And disabling copy-on-write also disables data checksumming for those files, because a checksum cannot be updated atomically with an in-place overwrite — you have given up the integrity guarantee that was a main reason to choose the filesystem. Note too that even a `nodatacow` file will copy a block once after a snapshot is taken, since the snapshot's reference must be preserved. There is also the `autodefrag` mount option, which rewrites small random writes more contiguously. It helps some workloads and hurts snapshot-heavy ones, because defragmenting rewrites extents and breaks the sharing that made the snapshots cheap — space usage can jump sharply. ## A snapshot is not a backup The last point is the one to say out loud. A snapshot shares blocks with the live data on the same device: it survives an `rm -rf` and a bad deployment, and it does not survive the device failing, the pool being destroyed, or the machine being lost. It is a fast local undo. It becomes a backup only once it has been replicated somewhere else — `btrfs send` of a read-only snapshot, or `zfs send`, into storage with an independent failure domain. ## What the interviewer wants That you can explain why a snapshot is instantaneous *and* why it can bankrupt a volume weeks later, which are the same fact seen from two ends, and that you know which workloads should not sit on copy-on-write at all.
- Someone deletes a 200 GB file on a snapshotted Btrfs volume and no space is returned. What has happened, and how do you actually reclaim it?The extents are still referenced by one or more snapshots, so removing the live directory entry frees nothing. You have to identify which snapshots pin the data — `btrfs subvolume list` plus `btrfs filesystem usage` for the real accounting — and delete those snapshots. Space is returned asynchronously as the cleaner thread drops the now-unreferenced extents.
- What does chattr +C do on Btrfs, and what are its two catches?It marks a file as no-copy-on-write, so overwrites happen in place and the file stops fragmenting. First catch: it only takes effect on a file with no data, so you set it on the parent directory and let new files inherit it. Second: it disables data checksumming for those files, and after a snapshot is taken each block is still copied once.
- Is a Btrfs or ZFS snapshot a backup?No. It shares blocks with the live data on the same device, so it protects against accidental deletion, a bad upgrade or a bad migration, and against nothing else. Lose the device or the pool and the snapshot goes with it. It becomes a backup only once it has been replicated to independent storage, for example by sending a read-only snapshot elsewhere.
- Why is plain df untrustworthy on Btrfs?Because df reports a single free-space number and Btrfs's accounting is not one number. Space can be shared between subvolumes and snapshots, metadata is allocated in its own chunks, and raw space allocated to chunks is not the same as space usable for new data — a filesystem can fail writes with unallocated raw space remaining. `btrfs filesystem usage` reports the breakdown that actually explains it.
saying these in an interview costs you the question
- Believes a snapshot copies the data when it is taken
- Treats a same-device snapshot as a backup
- Expects deleting files to free space while snapshots exist
- Reads df on Btrfs as an accurate free-space number
- Thinks nodatacow keeps checksumming for those files