Ch. 26 · Apache Kafka

Kafka Retention vs Log Compaction: How Old Data Is Removed

How Kafka deletes old segments by time or size, how log compaction keeps the latest value per key, and how tombstones delete keys.

~8 min readintermediateupdated Oct 6, 2026

Interviewers ask “how long does Kafka keep messages?” to see whether you understand that Kafka is a log, not a queue. The follow-ups go deeper: “what is log compaction?”, “how do you delete a key?”, “why can I still read records older than retention.ms?” and “what happens to a consumer whose offset was deleted?” These come up when a disk fills, when a privacy request requires deleting a customer’s data, or when a service that was down for a week restarts.

This guide explains how a partition is stored on disk, how time and size retention delete it, how compaction keeps the latest value per key, and the settings that control both. It targets Kafka 3.8/4.0. CLI commands are illustrative; the compaction model at the end is a small Python script that was run.

Before you start

You should know that a topic is split into partitions and that each record has an offset within its partition. Knowing what a record key is matters for compaction, because compaction works per key. Familiarity with the idea of an append-only file helps.

The short answer

Reading does not delete anything in Kafka. Each partition is stored as a series of segment files, and the topic’s cleanup.policy decides how old data goes. With delete (the default), closed segments are removed once their newest record is older than retention.ms (7 days by default) or the partition is larger than retention.bytes (unlimited by default). With compact, a background cleaner keeps at least the latest record for each key and drops older ones, and a record with a null value, a tombstone, deletes the key. Both work only on closed segments, never on the active one.

How it works

On disk, a partition is a directory such as orders-0/. Inside it, each segment is a set of files named after its first offset: 00000000000000000000.log holds the records, with .index and .timeindex files mapping offsets and timestamps to file positions. New records go to the newest, active segment. It is closed, or “rolled”, when it reaches segment.bytes (1 GiB) or is older than segment.ms (7 days), and a new active segment starts.

Time and size retention run periodically, every log.retention.check.interval.ms (5 minutes). A closed segment is eligible when its largest record timestamp is older than retention.ms, or when removing the oldest segments is needed to bring the partition under retention.bytes. Deletion is per segment: the broker renames the files with a .deleted suffix and removes them shortly afterwards. The partition’s log start offset moves forward; offsets are never reused.

Compaction is done by log cleaner threads. For a compacted partition, the cleaner reads the “dirty” part of the log (written since the last clean), builds a map from each key to its newest offset, then rewrites the older segments keeping only records that are still the newest for their key. A partition becomes eligible when the dirty fraction exceeds min.cleanable.dirty.ratio (0.5), or when a record has waited longer than max.compaction.lag.ms. Records younger than min.compaction.lag.ms are left alone.

Setting Default Applies to
cleanup.policy delete delete, compact or compact,delete
retention.ms 604800000 (7 days) delete policy
retention.bytes -1 (no limit) delete policy, per partition
segment.bytes / segment.ms 1 GiB / 7 days when the active segment rolls
delete.retention.ms 86400000 (24 hours) how long tombstones survive
min.compaction.lag.ms / max.compaction.lag.ms 0 / unbounded compaction timing

Compaction guarantees are precise and worth quoting: every consumer that keeps up sees every record, the latest value for each key survives, offsets never change, and order within the partition is preserved. It does not guarantee that only one value per key exists at any moment.

Step-by-step walkthrough

Step 1: Pick the policy from what the topic represents

An event stream (“payment captured”, “page viewed”) is a sequence of facts, so it uses delete with a retention that covers your replay needs. A changelog (“current email of user 7”, a Kafka Streams state store, connector offsets) represents current state, so it uses compact. Internal topics such as __consumer_offsets are compacted for the same reason.

kafka-topics.sh --bootstrap-server localhost:9092 --create --topic page-views \
  --partitions 12 --replication-factor 3 --config retention.ms=259200000      # 3 days

kafka-topics.sh --bootstrap-server localhost:9092 --create --topic user-profiles \
  --partitions 6 --replication-factor 3 --config cleanup.policy=compact
Terminal

compact,delete combines both: the latest value per key is kept, but even those are removed after the retention time. It suits keyed data you only need for a bounded window.

Step 2: Size retention for the slowest consumer and the disk

Retention must cover the longest outage any consumer group can have, plus time to notice and fix it. Then check the disk: bytes per day multiplied by retention days multiplied by the replication factor.

kafka-configs.sh --bootstrap-server localhost:9092 --alter --entity-type topics \
  --entity-name page-views --add-config retention.ms=604800000,retention.bytes=53687091200
Terminal

When both limits are set, whichever is reached first triggers deletion. For long retention on cheap storage, tiered storage (production-ready since Kafka 3.9) keeps recent segments on brokers for local.retention.ms and older ones in object storage, though it does not support compacted topics.

Step 3: Delete a key with a tombstone

On a compacted topic you delete a key by producing the key with a null value. Kafka rejects records with a null key on compacted topics, because they could never be compacted.

// Illustrative: delete user-7 from a compacted topic
producer.send(new ProducerRecord<String, String>("user-profiles", "user-7", null));
java

Consumers must treat a null value as “delete this key” in their own store. The cleaner keeps the tombstone for at least delete.retention.ms after compacting it, so a consumer that is behind still sees the delete, and then removes it too.

Step 4: Bound how long old values survive

Compaction is lazy. With the defaults, a low-traffic partition may not reach the dirty ratio for days, and records in the active segment are not touched at all. When a deadline matters, set max.compaction.lag.ms to bound how long a record can stay uncompacted, and a segment.ms no larger than that so the active segment rolls in time.

kafka-configs.sh --bootstrap-server localhost:9092 --alter --entity-type topics \
  --entity-name user-profiles \
  --add-config max.compaction.lag.ms=86400000,segment.ms=43200000,delete.retention.ms=3600000
Terminal

Smaller segments mean more files and more cleaning work, so tighten these only for topics with a real deadline.

Worked scenario

A customer asks for their account to be erased, and the profile service publishes a tombstone for user-7 to the compacted user-profiles topic. Two weeks later an audit replays the topic from the beginning and still finds user-7 with an old email address.

The topic used default settings and receives a few hundred updates a day. The active segment had not reached 1 GiB, and segment.ms is 7 days, so for a week the tombstone and the old values sat together in the active segment, which is never compacted. After the roll, the dirty ratio was still below 0.5, so the cleaner did not pick the partition. Nothing was broken; compaction makes no promise about timing.

The fix sets a deadline: max.compaction.lag.ms=86400000 and segment.ms=43200000, so any record becomes eligible for compaction within about a day and a half. The team also documents that erasure on Kafka takes up to that long, and that downstream stores built from the topic must apply the tombstone themselves.

A second incident involved retention: a reporting consumer was down for 9 days on a topic with 7-day retention. Its committed offset was below the log start offset, so it applied auto.offset.reset=latest and skipped two days of data that still existed. Alerting on time lag relative to the retention window catches this early.

Common mistake

  • “Consumed messages are deleted.” Reading has no effect on retention.
  • “retention.ms=1h means nothing older than an hour is readable.” Deletion is per closed segment, and the active segment can be days old.
  • “Compaction leaves exactly one record per key.” It keeps at least the latest; older values remain until the cleaner reaches them.
  • “Compaction renumbers offsets.” Offsets stay; consumers skip the gaps.
  • “A tombstone deletes the key immediately.” It is just another record until compaction runs, and the tombstone itself disappears after delete.retention.ms, so a consumer that is further behind than that can miss the delete.

Verify the behavior

This toy model in Python mirrors the cleaner’s rules on six records, with offset 5 still in the active segment. It printed the output in the comments when run with Python 3:

log = [(0, "user-1", "alice@old.example"), (1, "user-2", "bob@example.com"),
       (2, "user-1", "alice@new.example"), (3, "user-3", "carol@example.com"),
       (4, "user-3", None), (5, "user-2", "bob@work.example")]
active_from = 5

def compact(log, active_from, drop_tombstones):
    closed = [r for r in log if r[0] < active_from]
    active = [r for r in log if r[0] >= active_from]
    latest = {key: offset for offset, key, _ in closed}   # only up to the active segment
    kept = [r for r in closed if latest[r[1]] == r[0] and not (drop_tombstones and r[2] is None)]
    return kept + active

print([r[0] for r in compact(log, active_from, False)])   # [1, 2, 4, 5]
print([r[0] for r in compact(log, active_from, True)])    # [1, 2, 5]
python

Offset 1 survives because its newer value is in the active segment, and offsets keep their gaps. On a real broker, inspect segments and settings directly:

kafka-configs.sh --bootstrap-server localhost:9092 --describe --entity-type topics --entity-name user-profiles
kafka-log-dirs.sh --bootstrap-server localhost:9092 --describe --topic-list user-profiles
kafka-dump-log.sh --files /var/lib/kafka/data/user-profiles-0/00000000000000000000.log --print-data-log
Terminal

kafka-dump-log.sh prints each record’s offset, key and payload, so you can see which records the cleaner removed.

Follow-up questions

What happens to a consumer whose committed offset was deleted? Its fetch is out of range and it falls back to auto.offset.reset, skipping to the end with latest or replaying everything with earliest.

Can you delete specific records from a delete-policy topic? Only a prefix: kafka-delete-records.sh advances the log start offset of a partition. There is no per-record delete.

Why do Kafka Streams state stores use compacted changelogs? A restarted instance rebuilds its store by replaying the changelog, and compaction keeps that replay proportional to the number of keys, not the number of updates.

Interview exercise

A compacted topic contains, in order, k1=A (offset 0), k2=B (1), k1=C (2), k2=null (3) and k1=D (4). Offsets 0 to 3 are in closed segments and offset 4 is in the active segment. What can a new consumer reading from the beginning see after one compaction pass, and after delete.retention.ms has also passed?

Answer and reasoning

The cleaner only maps keys in closed segments, so within offsets 0 to 3 the newest k1 is offset 2 and the newest k2 is the tombstone at offset 3. After one pass the consumer sees offsets 2 (k1=C), 3 (k2=null) and 4 (k1=D). Offset 2 survives because its replacement, offset 4, is still in the active segment. Once delete.retention.ms has passed, a later pass drops the tombstone, leaving offsets 2 and 4. When offset 4’s segment rolls and is cleaned, offset 2 disappears too. The consumer ends with k1=D and no k2 in every case, which is the guarantee compaction actually gives.

Continue learning

Practise with the Apache Kafka interview questions and the Kafka MCQs. Related notes: Kafka partitions, offsets and ordering and event sourcing in microservices. Primary sources: the Apache Kafka documentation, Confluent’s log compaction design notes and the topic configuration reference.

More in Apache Kafka

esc