Kafka

Kafka stores streams of events in topics split into partitions — append-only logs. Producers write events; consumer groups read them in parallel and remember their position with offsets.

Topic "user-events", 3 partitions

Producer

Ready to send events.

Topic "user-events"

  1. P0
  2. P1
  3. P2

Group "billing"

Not started yet.

A topic is split into partitions — append-only logs that can live on different brokers and be read in parallel.

Step 1 / 27
Produced
0
Consumed
0
Lag
0
Consumers3 partitions — try 4 consumers.

What's happening?

  1. The producer hashes each event's key to pick a partition, so all events for one key stay in order in one partition.
  2. A consumer group gives each partition to exactly one of its consumers; with more consumers than partitions, some sit idle.
  3. Consumers commit offsets as they go. If one crashes, the group rebalances and its partitions resume from the last committed offset.

Complexity

Time
Appends and sequential reads are O(1) per event
Space
Events are kept for the retention period, read or not

Where you'll meet it

Order and payment events, activity tracking, log pipelines, and letting many services react to the same event without calling each other.

Common mistake

Adding consumers to go faster without adding partitions. Parallelism is capped at the partition count — extra consumers just wait.

FAQ

Why do events need a key?

The key picks the partition. Same key, same partition — which is what keeps one user's or one order's events in order.

What is consumer lag?

How many events have been written but not yet processed by a group. Growing lag means consumers cannot keep up.

Is Kafka a queue?

Not exactly: reading an event does not delete it. Each consumer group keeps its own offsets, so many groups can read the same events.

Is this how Kafka hashes keys?

No — real Kafka uses the murmur2 hash. A character sum keeps the example easy to follow.