On an ext4 filesystem, what does the journal actually protect after an unclean shutdown, and how do the data=ordered, data=writeback and data=journal mount options differ?
answer
- metadata, not your bytes
- ordering between data and metadata
- stale blocks in a freshly grown file
- journal it twice, pay twice
- fsync is still the app's job
basics
~20 sext4's journal protects filesystem metadata consistency, not file contents. The default data=ordered flushes data blocks before committing the metadata that points at them; data=writeback drops that ordering and can expose stale bytes; data=journal journals data too, more slowly.
solid answer
~50 sext4 journals metadata — inodes, bitmaps, directory blocks, extent trees — through the jbd2 layer, so after a crash the kernel replays the journal at mount time and the structure is consistent in seconds instead of hours of full checking. What the journal does *not* promise is that your file's contents are correct. That is what the `data=` mode governs. `data=ordered`, the default, writes a file's data blocks out before committing the metadata that references them, so a file never ends up pointing at blocks it does not own. `data=writeback` journals metadata only with no ordering against data, which is faster but can leave a newly extended file exposing whatever was previously in those blocks. `data=journal` puts file data through the journal as well — the strongest, and roughly the slowest, because everything is written twice. None of them substitutes for `fsync()` in the application.
code
bash · 2 linestune2fs -l /dev/sdb1 | grep -E 'Filesystem features|Default mount options'
findmnt -no OPTIONS /srvgo deeper
Know that ext4 keeps a journal so the filesystem comes back consistent quickly after a crash, and that data=ordered is the default. Say clearly that the journal is about filesystem structure, not about your file's contents.
Explain the three data modes and the ordering each one does or does not impose between file data and the metadata that points at it, including how writeback can leave stale block contents in a newly extended file.
Show the judgment call: which volumes can accept writeback, why data=journal's double write is usually the wrong trade, and how you would argue that a crash-consistency complaint is really a missing fsync in the application.
Own the durability story end to end — filesystem mode, application fsync discipline, and whether the storage layer honours cache flush requests at all. Be ready to say what the organisation's data-loss window actually is and who is accountable for it.
## Why a journal exists at all A single logical operation — appending to a file — touches several independent structures on disk: the block allocation bitmap, the inode's size and extent tree, possibly a directory entry. If power is lost between those writes, the filesystem is internally inconsistent: blocks marked used that no file claims, or worse, a file whose extents point at blocks the allocator also handed to someone else. Historically the fix was a full check at boot (`fsck`), whose cost scales with the size of the filesystem, and which on a multi-terabyte volume means a very long outage. A journal turns that into a bounded operation. ext4 writes a description of the change to a dedicated on-disk area first, commits it, and only then applies the change in place. After a crash the kernel finds the journal non-empty at mount time and replays it, and the filesystem is consistent again. This is why an ext4 root filesystem comes back in seconds after an unclean reboot rather than sitting at a progress bar. ## Metadata is journaled; data is a separate question The crucial distinction is that, by default, only metadata goes through the journal. A consistent filesystem is not the same as a correct file. Consider appending 4 KiB to a file: the metadata change says "this inode now owns block N and is 4096 bytes longer". If that metadata commit lands but the 4 KiB of data never reaches block N, the filesystem is perfectly consistent — and your file now contains whatever bytes block N held previously, possibly from a deleted file belonging to someone else. That is both a correctness problem and an information-disclosure problem. The `data=` mount option is the knob that decides how much the filesystem does about this. ## The three modes **`data=ordered` (the default).** Metadata is journaled, and ext4 additionally guarantees ordering: a transaction's data blocks are pushed to the device *before* the metadata transaction that references them is committed. You never see metadata pointing at blocks whose contents were never written, so the stale-block exposure above cannot happen. The cost is a write barrier between data and metadata, which serialises some workloads. **`data=writeback`.** Metadata is journaled, but there is no ordering constraint against data. Metadata may be committed first, and if the machine dies in the window before the data lands, the file legitimately contains stale block contents. It is the fastest of the three and is defensible for filesystems where you genuinely do not care about the contents of files being written at the moment of a crash — a scratch or cache volume — and indefensible for a general-purpose or multi-tenant filesystem. **`data=journal`.** File data is written into the journal as well as to its final location. Nothing can be torn or stale, at the price of writing every byte twice; throughput on write-heavy workloads falls sharply. It also disables ext4's delayed allocation. It is used where correctness dominates and the write volume is modest. The mode is chosen at mount time. You make it persistent in the fstab entry or by setting the filesystem's default with `tune2fs`, and you cannot switch modes on a mounted filesystem: ``` tune2fs -l /dev/sdb1 | grep -E 'Filesystem features|Default mount options' tune2fs -o journal_data_writeback /dev/sdb1 ``` ## Delayed allocation, and the famous zero-length files Modern ext4 uses extents (contiguous block ranges recorded as start-plus-length) instead of ext3's indirect block maps, and it uses *delayed allocation*: when you write, the data sits in the page cache and no blocks are chosen until writeback, which lets the allocator pick better, more contiguous placement. That is a real performance win, and it lengthened the window in which a crash loses recent writes. When ext4 became widespread, users reported files that were zero length after a crash — applications had used the write-a-temporary-file-then-`rename()`-over-the-original idiom without calling `fsync()` before the rename, so the rename's metadata was durable while the data was still only in the page cache. ext4 gained heuristics that force allocation and writeback for these truncate-and-replace and rename-replace patterns, but the underlying lesson stands: the filesystem's ordering guarantees are about its own structures, and durability of your bytes is what `fsync()` (or `fdatasync()`) is for. A database that omits it is not saved by any `data=` mode. ## What to say in the interview Separate the two guarantees explicitly. The journal buys fast, bounded recovery of the filesystem's structure. `data=ordered` additionally buys the promise that a file never exposes blocks whose contents were not written. Neither buys durability of a specific `write()` — only `fsync()` does, and only if the device is not lying about its own cache flushes.
- With data=ordered in force, is my data on disk once write() returns?No. `data=ordered` only constrains the order in which the filesystem commits data relative to the metadata that references it; the data may still sit in the page cache for seconds. Durability for a specific write requires `fsync()` or `fdatasync()` on the file, plus `fsync()` on the directory if you also created or renamed the entry — and a device that honours cache flushes.
- When would you actually accept data=writeback on a production volume?On a filesystem whose contents are disposable and rebuilt after any crash — a build cache, a scratch or spool area, a rendering workspace — where the only thing that matters is that the filesystem itself mounts clean. Never on a shared or multi-tenant volume, because a crash can leave freshly extended files exposing block contents that belonged to something else.
- What is delayed allocation, and why did it produce reports of empty files after crashes?Delayed allocation keeps written data in the page cache and defers choosing blocks until writeback, which yields better contiguity. It also widened the crash window, so applications using write-temp-then-rename without an intervening `fsync()` could end up with a durable rename over a file whose data was never written. ext4 added heuristics for the rename and truncate-replace patterns, but `fsync()` remains the real fix.
- Can you disable the ext4 journal entirely, and would you?Yes — the `has_journal` feature can be cleared with `tune2fs` on an unmounted filesystem, and it is occasionally done for throwaway volumes on very fast devices. The cost is that any unclean shutdown now requires a full `e2fsck` whose duration scales with the filesystem, so you trade a small steady-state gain for an unbounded recovery time. Rarely worth it.
saying these in an interview costs you the question
- Thinks the journal guarantees your file contents survive a crash
- Believes data=ordered removes the need to call fsync()
- Says data=journal is the default because it is safest
- Confuses journal replay at mount with a full fsck run
- Thinks the journal is what makes ext4 use extents