Ch. 21 · Operating Systems

Operating Systems: Journaling and Durability

How write-ahead journaling keeps a filesystem consistent after a crash, and why durability needs an explicit fsync.

~2 min readadvancedupdated Oct 5, 2026

A crash can interrupt a multi-step update, leaving on-disk structures half-written. Journaling writes the intended changes to a log first and applies them atomically, so recovery can complete or roll back. Durability, however, is separate: data may sit in a cache until it is explicitly flushed.

Before you start

You should understand filesystems and basic I/O. This article covers journaling and the durability guarantee.

Step-by-step walkthrough

Step 1: Write the journal before the data

A write-ahead journal records the metadata (and optionally data) for a change before applying it. If a crash happens mid-apply, recovery replays the journal to finish the operation, so the filesystem is never left in an inconsistent intermediate state.

Step 2: fsync makes a write durable

A write may return while data is still in the page cache, not on stable storage. fsync forces the data and metadata to disk, so it survives a power loss. Without fsync, an acknowledged write can be lost, which is why databases call it on commit.

Step 3: Understand journaling modes

Metadata-only journaling keeps structures consistent but may lose recently written file data. Data journaling writes data to the journal too, which is slower but safer. This is a trade-off between throughput and the durability window, chosen per workload.

Worked scenario

The sequence makes a file change durable.

int fd = open("data", O_WRONLY | O_CREAT, 0644);
write(fd, buf, len);   /* may still be in cache */
fsync(fd);             /* now on stable storage */
close(fd);
c

Walk through the example

write may complete while the bytes are cached, so a crash right after could lose them. fsync flushes them to disk before returning, which is the durability point. A database makes the same call before reporting a committed transaction.

Common mistake

Assuming write returning means the data is on disk, which loses acknowledged writes on a crash. Another is calling fsync on the file but not its directory, so a new file’s name may not be durable even though its data is.

Verify the behavior

Write and read back after fsync to confirm persistence. Simulate a crash before fsync and confirm the data can be lost. Compare throughput with and without fsync to see the cost of durability.

Interview exercise

Why must a database call fsync on commit?

Answer and reasoning

Because it must not acknowledge a transaction that a power loss could erase. write may leave data in the OS cache, so fsync forces it to stable storage before the commit is reported. This is the durability letter of ACID: a committed transaction survives a crash.

Continue learning

Compare durability in Filesystem durability and page cache behavior in Page cache. Read the Linux fsync documentation and try the Operating systems interview questions.

More in Operating Systems

read ✓Operating Systems · hard

Filesystem Writes, Buffers and Durability

Filesystem Writes, Buffers and Durability. Learn the reasoning, a practical example, common mistakes and an interview exercise.

~2 min readread →
read ✓Operating Systems · hard

Operating Systems: CPU Cache Locality

Improve performance with spatial and temporal locality, avoid pointer chasing, and understand false sharing between threads.

~2 min readread →
read ✓Operating Systems · hard

Copy-on-Write Memory and fork

How copy-on-write lets fork share pages until a write, why RSS can be misleading, and the implications for memory and latency.

~2 min readread →
esc