skip to content

You are provisioning two volumes on a Linux server: one for a busy relational database's data directory that may have to grow later, and one for a CI build cache holding millions of small files that is wiped weekly. How would you choose between ext4, XFS and Btrfs for each, and which property of each filesystem drives the choice?

level: seniorimportance: must knowfreq 56%

answer

  1. copy-on-write versus overwrite in place
  2. which direction can this volume resize
  3. who allocates inodes, and when
  4. millions of small files is a constraint
  5. alignment is a mkfs-time decision

basics

~20 s

Pick XFS or ext4 for the database — both journal metadata and overwrite in place, and XFS grows online but can never shrink. Btrfs suits the disposable build cache, where cheap snapshots, compression and checksums outweigh its copy-on-write fragmentation.

solid answer

~50 s

For the database volume the deciding properties are in-place overwrite and predictable growth. XFS and ext4 both update data blocks in place, so a random-write workload does not fragment the way a copy-on-write filesystem does; XFS's allocation groups give it an edge under heavy parallel I/O, and it is the default on RHEL. The catch is direction of growth: XFS grows online with `xfs_growfs` and can never be shrunk, whereas ext4 grows online and can shrink offline with `resize2fs`, so if the size estimate is shaky ext4 buys you an exit. For the build cache the workload is millions of small files with a weekly reset, so durability barely matters: XFS allocates inodes dynamically while ext4 fixes the inode count at mkfs time, and Btrfs is attractive if you want transparent compression and instant snapshot-and-rollback of the cache. I would also align the filesystem to the underlying RAID stripe at mkfs time in every case.

code

bash · 4 lines
bash
mkfs.xfs -f -d su=64k,sw=8 /dev/mapper/vg-data
mount /dev/mapper/vg-data /srv/data
lvextend -L +200G /dev/mapper/vg-data
xfs_growfs /srv/data

go deeper

for a junior

Know that ext4 is the traditional default, XFS is the default on RHEL, and Btrfs is the copy-on-write one with snapshots. Be able to say that filesystem choice is made when the volume is formatted.

for a middle

Explain the mechanics behind the choice: in-place overwrite versus copy-on-write, XFS allocation groups and dynamic inodes versus ext4's fixed inode count, and which filesystems can grow or shrink online.

for a senior

Demonstrate judgment on the irreversible parts — that XFS never shrinks, that the inode ratio and stripe alignment are set once at mkfs time — and tie the recommendation to the workload's actual write pattern rather than a feature list.

for a principal

Own the standardisation question: how many filesystems the organisation can operate competently, whether a second one earns its training and tooling cost, and how provisioning defaults keep teams from making one-way sizing decisions by accident.

## The three bets ext4, XFS and Btrfs are not three implementations of the same idea. They make different structural bets, and the interview question is whether you can name the bet rather than recite a feature matrix. **ext4** is the conservative journaling filesystem. It updates data in place, journals metadata, and uses extents — contiguous ranges recorded as start-plus-length — instead of the older indirect block maps. Its distinguishing operational property is maturity: `e2fsck` is the most battle-tested repair tool in the Linux world, and ext4 is the only one of the three that can be made *smaller*, offline, with `resize2fs`. **XFS** is the scalability bet, originating on SGI's IRIX for large parallel workloads. It divides the filesystem into allocation groups, each with its own free-space and inode structures, so independent threads can allocate concurrently without contending on one global structure. It allocates inodes dynamically rather than fixing a count at format time, uses delayed allocation aggressively, journals metadata only (there is no data-journalling mode), and protects its own structures with metadata checksums in the v5 on-disk format. It grows online and it never shrinks — not "shrinking is discouraged", but the operation does not exist. It is the default on RHEL and its derivatives. **Btrfs** is the copy-on-write bet. Nothing is overwritten in place: a modified block is written elsewhere and the tree pointing at it is updated. That structure is what buys snapshots that cost nothing to take, subvolumes, transparent compression, data checksums (not just metadata) and reflink copies. The same structure is what makes a random-overwrite workload fragment, and it is why free-space accounting is not a single number you can read from `df`. It is the default on Fedora Workstation and openSUSE. ZFS makes broadly the same bet with a more integrated volume manager, but it is not in the mainline kernel and ships out of tree for licensing reasons, which is a deployment consideration of its own. ## The database volume A relational database's data directory is dominated by random overwrites of fixed-size pages, plus a sequentially appended write-ahead log, and it usually issues its own `fsync()` calls and does its own crash recovery. What it wants from a filesystem is that a page rewritten in place stays where it was. On a copy-on-write filesystem it does not: each overwrite lands in a new location, so a table that was laid out contiguously becomes progressively scattered, and sequential scans degrade into random reads. Btrfs offers mitigations — the `nodatacow` mount option, or `chattr +C` applied to the directory before any files are created — but disabling copy-on-write also disables data checksumming for those files, which removes the main reason you chose Btrfs. So the realistic choice is ext4 or XFS, and I would take XFS for a busy multi-threaded server because allocation-group parallelism and its allocator behave well under concurrent load, unless the organisation is standardised on ext4 and nobody wants a second filesystem to operate. The growth clause in the question is the real trap. Both grow online, but only ext4 can go back: ``` xfs_growfs /srv/data # XFS: online grow only, no shrink at all resize2fs /dev/mapper/vg-data # ext4: grows online, shrinks only unmounted ``` If you cannot rule out over-provisioning, XFS commits you to a copy-and-swap migration to get the space back. That single asymmetry decides plenty of real arguments. ## The build cache volume Here the priorities invert. The data is disposable and regenerated, the file count is enormous and the files are small, and the volume is wiped weekly. The file count is the first constraint. ext4 fixes its inode count when the filesystem is created, from the bytes-per-inode ratio; a filesystem full of tiny files can therefore run out of inodes while reporting free space, and the only fix is to re-create it with a different ratio. XFS allocates inodes on demand and does not have that failure mode, which alone makes it a good default for this shape of workload. Btrfs is the interesting alternative, because what a build cache actually wants is cheap resets and cheap space: transparent compression on highly compressible artefacts, and a snapshot you take before a run and roll back to afterwards in constant time. The copy-on-write fragmentation that disqualified it for the database costs little here, because nothing is randomly overwritten for months at a stretch — files are written once, read, and eventually deleted wholesale. The price is operational surface: more knobs, more distinct failure modes, and a repair story that is less well trodden than `e2fsck`. ## The thing everyone forgets: alignment Whatever you pick, the filesystem should be told about the geometry beneath it. On RAID or a striped device, `mkfs.xfs` takes stripe unit and stripe width (`-d su=,sw=`) and `mke2fs` takes `-E stride=,stripe-width=`, so allocations line up with the stripe instead of straddling it and forcing read-modify-write. Modern tooling often detects this through device-mapper, but on hardware RAID it frequently cannot, and a misaligned filesystem quietly loses a large fraction of the array's write throughput for its entire life. It is set once, at format time, and cannot be fixed later without re-creating the filesystem. ## How to answer Do not deliver a feature table. Name the workload property that decides — overwrite pattern, growth direction, file count, disposability — and let the filesystem fall out of it. Then say which decision is irreversible, because that is the part a senior engineer is being paid to notice.

  • The database volume was built on XFS at 4 TB and the team now wants it at 1 TB. What do you tell them?
    That XFS has no shrink operation at all, so there is nothing to run. The only path is to create a smaller filesystem on new space, copy the data across with the service stopped or with a replica taking traffic, and swap the mount. If the volume sits on LVM the copy is easier to stage, but you still cannot reduce the XFS filesystem itself.
  • How would you decide between putting the build cache on XFS versus Btrfs?
    Ask whether anyone will actually use the copy-on-write features. If the workflow benefits from snapshot-and-rollback between runs, or the artefacts compress well, Btrfs earns its extra operational surface. If the cache is just a large pile of files that gets wiped, XFS gives you dynamic inode allocation and a simpler failure model, and I would take the simpler thing.
  • When does the ext4 inode count actually become a problem, and can it be fixed after the fact?
    When the average file is far smaller than the bytes-per-inode ratio chosen at format time — build trees, mail spools, caches. The filesystem then reports free space while creation fails with ENOSPC because inodes are exhausted. It cannot be raised on an existing filesystem: you re-create it with a denser ratio via mke2fs -i or -N and restore the data.
  • Does the choice change if the volume is a cloud block device rather than local disks?
    The alignment argument weakens, since the provider hides the geometry, but the resize asymmetry gets sharper: cloud volumes are easy to grow and usually impossible to shrink, so an XFS filesystem on one is doubly one-way. The overwrite-pattern argument for keeping a database off copy-on-write is unchanged.

saying these in an interview costs you the question

  • Claims XFS can be shrunk with the right tool
  • Recommends Btrfs for a write-heavy database without qualification
  • Treats the three as interchangeable because 'they're all journaling'
  • Ignores inode exhaustion for millions of small files
  • Never mentions that alignment is fixed at format time

context