“Does Kafka guarantee message order?” is usually the first Kafka question in a backend interview, and the honest answer starts with “it depends on the partition.” Interviewers follow up with “how would you keep all events for one order in sequence?”, “what is an offset?” and “what happens when you add partitions?” They ask because ordering bugs are some of the most expensive Kafka mistakes: a PAID event processed before CREATED, or an account balance computed from deposits in the wrong order.
This guide explains topics, partitions and offsets from the ground up, then shows how keys give you per-entity ordering and exactly where that guarantee breaks. It targets Kafka 3.8/4.0 in KRaft mode and the Java client. Java and CLI snippets are illustrative and were not run against a live cluster; the partition numbers come from a tested port of the client’s hash function.
Before you start
You should know what a producer and a consumer are in general terms, and be comfortable reading basic Java. It helps to know what a hash function is and what “modulo” means, because that is the entire partitioning algorithm. No Kafka internals knowledge is assumed.
The short answer
Kafka guarantees order within a partition, not across a topic. Each partition is an append-only log, and every record in it gets an increasing offset; consumers read a partition in offset order. To keep related events in sequence, give them the same key: the producer hashes the key to pick the partition, so all events for order-1 land in the same partition and are read in the order they were written. That holds as long as the partition count does not change and the producer does not reorder on retry, which the default idempotent producer prevents.
How it works
A topic such as orders is a logical name. Physically it is split into partitions, orders-0 through orders-5 for a six-partition topic. Each partition is an independent log stored on one leader broker and copied to followers. Records are only ever appended to the end of a partition, and each one gets the next offset: 0, 1, 2 and so on.
Offsets are local to a partition. Offset 42 in orders-0 and offset 42 in orders-3 are unrelated records, and nothing links their timing. That is why there is no topic-wide order: a consumer reading partitions 0 and 3 receives batches from each as they arrive, interleaved however the network and the fetch sizes decide.
The producer decides the partition for every record, in this order:
- If the
ProducerRecordnames a partition, that partition is used. - If the record has a key, the partition is
toPositive(murmur2(keyBytes)) % numPartitions, computed on the serialized key bytes. - With no key, the sticky partitioner fills a batch for one partition, then moves to another. Since Kafka 3.3 it also sends less data to brokers that respond slowly.
// Illustrative: three ways to choose a partition
producer.send(new ProducerRecord<>("orders", 2, "order-1", json)); // explicit partition 2
producer.send(new ProducerRecord<>("orders", "order-1", json)); // keyed: hash of "order-1"
producer.send(new ProducerRecord<>("orders", json)); // no key: sticky partitionerThere are several offsets worth naming, because interviewers like to test whether you mix them up:
- Log end offset (LEO): the offset the next appended record will get.
- High watermark: the highest offset replicated to all in-sync replicas. Consumers can only read below it.
- Position: the offset of the next record this consumer will fetch, held in memory.
- Committed offset: the position a consumer group saved to the
__consumer_offsetstopic. It is the next record to read after a restart, not the last one processed.
Offsets are never reused within a partition, but they can have gaps. Compaction removes records without renumbering, and transaction markers occupy offsets that consumers never see.
Step-by-step walkthrough
Step 1: Create a topic and decide the partition count
The partition count is a capacity decision. It caps how many consumers in one group can work in parallel, and for keyed topics it should be chosen with growth in mind, because changing it later remaps keys.
kafka-topics.sh --bootstrap-server localhost:9092 --create --topic orders \
--partitions 6 --replication-factor 3
kafka-topics.sh --bootstrap-server localhost:9092 --describe --topic orders
# One line per partition: Partition, Leader, Replicas, Isr--describe shows one line per partition with its leader broker and replicas, which makes it concrete that a topic is really six independent logs.
Step 2: Key every event by the entity whose order matters
Ask “which events must be processed in sequence relative to each other?” For an order lifecycle the answer is all events of the same order, so the key is the order ID. Events for different orders are independent and can be processed in parallel.
// Illustrative
String key = event.orderId(); // "order-1"
producer.send(new ProducerRecord<>("orders", key, toJson(event)), (meta, ex) -> {
if (ex != null) {
log.error("send failed for {}", key, ex);
} else {
log.debug("{} -> partition {} offset {}", key, meta.partition(), meta.offset());
}
});The callback reports the partition and offset the broker assigned, which is the easiest way to confirm the mapping in logs. Picking the key is the real design decision: key by customer ID and every order from one customer is serialized behind the others; key by a random UUID and you have no ordering at all.
Step 3: Keep the producer from reordering
The producer pipelines requests: up to max.in.flight.requests.per.connection (5 by default) batches can be in flight to a broker. Without idempotence, if batch 1 fails and is retried after batch 2 succeeded, the partition ends up with batch 2 before batch 1. The idempotent producer, on by default since Kafka 3.0, numbers batches per partition so the broker rejects out-of-order writes and the producer resends in sequence.
props.put(ProducerConfig.ENABLE_IDEMPOTENCE_CONFIG, true); // default, but explicit is clearer
props.put(ProducerConfig.ACKS_CONFIG, "all"); // required by idempotenceSetting it explicitly matters because conflicting settings such as acks=1 quietly turn the default off.
Step 4: Keep the consumer from reordering
Kafka delivers a partition in order, but your code can still scramble it. Submitting each record to a thread pool, or sending failed records to a retry topic while later ones carry on, breaks per-key order. If you need parallelism inside a consumer, partition the work by key: hash the key to one of N single-threaded workers, so each key is still handled by one thread in sequence.
// Illustrative: per-key parallelism without losing per-key order
ExecutorService[] lanes = new ExecutorService[8];
for (int i = 0; i < lanes.length; i++) lanes[i] = Executors.newSingleThreadExecutor();
for (ConsumerRecord<String, String> r : records) {
int lane = (r.key().hashCode() & 0x7fffffff) % lanes.length;
lanes[lane].submit(() -> handle(r));
}This keeps order per key but makes committing harder, because an offset can only be committed once every earlier record in that partition is done.
Worked scenario
A payments team publishes PaymentAuthorized and PaymentCaptured events. The producer was written without a key, so the sticky partitioner spreads events across partitions. Under load, a consumer reading partition 1 processes PaymentCaptured for payment p-77 while PaymentAuthorized for the same payment is still waiting in partition 4. The capture handler finds no authorization and rejects it.
// Broken: no key, so related events can land in different partitions
producer.send(new ProducerRecord<>("payments", toJson(event)));The fix is to key by the payment ID so both events share a partition, and to make the handler tolerate replays:
// Fixed: same payment, same partition, same order
producer.send(new ProducerRecord<>("payments", event.paymentId(), toJson(event)));A month later the team raises the partition count from 6 to 8 for throughput. For a few minutes after the change, some payments again arrive out of order. Running the client’s hash on sample keys (see the verification section) shows why: the key order-1, for example, maps to partition 4 with 6 partitions and to partition 6 with 8. Old events for such a key sit in partition 4, new ones go to partition 6, and two consumers race. The safe options are to pause producers, let consumers drain the existing partitions, add partitions and then resume, or to create a new topic with the larger count and migrate consumers to it.
Common mistake
- “Kafka guarantees order in a topic.” Only per partition. A topic with one partition is fully ordered, but it also has a parallelism of one.
- “The committed offset is the last record processed.” It is the next record to read. Committing
record.offset()instead ofrecord.offset() + 1replays one record per partition on every restart. - “Offsets are contiguous, so
end - startis the record count.” Compaction and transaction markers leave gaps. - “Adding partitions is a harmless scaling step.” For keyed topics it remaps keys and breaks ordering across the change. Partitions can never be removed.
- “More consumers always means more throughput.” Within a group, consumers beyond the partition count sit idle.
Verify the behavior
The key-to-partition mapping is deterministic, so you can reproduce it outside Kafka. This JavaScript port of the client’s murmur2 and toPositive matched Kafka’s published test vectors (for example murmur2("21") = -973932308) and produced this output when run with Node 22:
// partition-demo.mjs, using a port of org.apache.kafka.common.utils.Utils.murmur2
const partitionFor = (key, n) => (murmur2(Buffer.from(key, 'utf8')) & 0x7fffffff) % n;
for (const key of ['order-1', 'order-2', 'order-3', 'order-4', 'customer-42']) {
console.log(key.padEnd(12), '6 partitions ->', partitionFor(key, 6), '| 8 partitions ->', partitionFor(key, 8));
}
// order-1 6 partitions -> 4 | 8 partitions -> 6
// order-2 6 partitions -> 3 | 8 partitions -> 3
// order-3 6 partitions -> 3 | 8 partitions -> 7
// order-4 6 partitions -> 2 | 8 partitions -> 2
// customer-42 6 partitions -> 3 | 8 partitions -> 1On a real cluster, produce keyed records and read them back with the partition and offset printed:
kafka-console-producer.sh --bootstrap-server localhost:9092 --topic orders \
--property parse.key=true --property key.separator=:
# order-1:CREATED
# order-1:PAID
kafka-console-consumer.sh --bootstrap-server localhost:9092 --topic orders --from-beginning \
--property print.key=true --property print.partition=true --property print.offset=trueBoth order-1 lines should show the same partition and increasing offsets. kafka-get-offsets.sh --topic orders prints the log end offset per partition.
Follow-up questions
How would you get a total order across a whole topic? Use one partition. You give up parallelism and put the whole stream on one broker, so it only suits low-volume streams.
Can two records in different partitions have the same offset? Yes. Offsets are per partition, so a record is identified by topic, partition and offset together.
What causes a hot partition? A skewed key: one huge tenant, a default value such as "unknown", or a low-cardinality key like a country code. Salting the key spreads load but gives up ordering for that key.
Interview exercise
A topic has 6 partitions. A producer with default settings sends order-1:CREATED, order-2:CREATED, order-1:PAID. The partition count is then raised to 8, and the producer sends order-1:SHIPPED. Using the mapping above, where does each record go, and which ordering guarantees still hold for a consumer group reading the topic?
Answer and reasoning
With 6 partitions, both order-1 records go to partition 4 and order-2:CREATED goes to partition 3. After the change, order-1:SHIPPED goes to partition 6, because the hash is taken modulo the new count. CREATED before PAID is still guaranteed, since they share partition 4. Nothing guarantees that SHIPPED is processed after PAID: they are in different partitions, possibly owned by different consumers, and the one owning partition 6 may be faster. There is never any ordering between order-1 and order-2. A strong answer names the fix too: drain the old partitions before resuming producers, or migrate to a new topic sized with headroom.
Continue learning
Practise with the Apache Kafka interview questions and the Kafka MCQs. Related notes: Kafka consumer groups and rebalancing, event ordering in microservices and partitioning in system design. Primary sources: the Apache Kafka documentation, the KafkaProducer Javadoc for Kafka 4.0 and Confluent’s Kafka producer design notes.