A crash can interrupt a multi-step update, leaving on-disk structures half-written. Journaling writes the intended changes to a log first and applies them atomically, so recovery can complete or roll back. Durability, however, is separate: data may sit in a cache until it is explicitly flushed.
Before you start
You should understand filesystems and basic I/O. This article covers journaling and the durability guarantee.
Step-by-step walkthrough
Step 1: Write the journal before the data
A write-ahead journal records the metadata (and optionally data) for a change before applying it. If a crash happens mid-apply, recovery replays the journal to finish the operation, so the filesystem is never left in an inconsistent intermediate state.
Step 2: fsync makes a write durable
A write may return while data is still in the page cache, not on stable storage. fsync forces the data and metadata to disk, so it survives a power loss. Without fsync, an acknowledged write can be lost, which is why databases call it on commit.
Step 3: Understand journaling modes
Metadata-only journaling keeps structures consistent but may lose recently written file data. Data journaling writes data to the journal too, which is slower but safer. This is a trade-off between throughput and the durability window, chosen per workload.
Worked scenario
The sequence makes a file change durable.
int fd = open("data", O_WRONLY | O_CREAT, 0644);
write(fd, buf, len); /* may still be in cache */
fsync(fd); /* now on stable storage */
close(fd);Walk through the example
write may complete while the bytes are cached, so a crash right after could lose them. fsync flushes them to disk before returning, which is the durability point. A database makes the same call before reporting a committed transaction.
Common mistake
Assuming write returning means the data is on disk, which loses acknowledged writes on a crash. Another is calling fsync on the file but not its directory, so a new file’s name may not be durable even though its data is.
Verify the behavior
Write and read back after fsync to confirm persistence. Simulate a crash before fsync and confirm the data can be lost. Compare throughput with and without fsync to see the cost of durability.
Interview exercise
Why must a database call fsync on commit?
Answer and reasoning
Because it must not acknowledge a transaction that a power loss could erase. write may leave data in the OS cache, so fsync forces it to stable storage before the commit is reported. This is the durability letter of ACID: a committed transaction survives a crash.
Continue learning
Compare durability in Filesystem durability and page cache behavior in Page cache. Read the Linux fsync documentation and try the Operating systems interview questions.