Event-Driven Architecture with Kafka: Topic Design, Partitioning, and Idempotency

The failure modes of streaming systems are quiet: duplicates, hot partitions, lagging consumers. Design decisions that prevent them, from partition keys to replay-safe consumers.

Event-Driven Architecture with Kafka: Topic Design, Partitioning, and Idempotency
Tomislav Pree
July 14, 2026
/
Tutorials

Topics are contracts

I treat every Kafka topic as a contract between whoever produces to it and whoever consumes from it, not just a name in a config file. Naming conventions matter. A topic's name should tell you what domain event it carries and who owns it, not just what service happens to write to it today. Schema evolution needs to be an explicit, versioned decision rather than something that happens implicitly because a producer added a field. A schema registry with compatibility rules enforced at write time catches that before it becomes a consumer's runtime exception. Retention is the third decision I make deliberately for every topic: how long events need to live, and whether this topic is a transient message bus or a durable log that downstream systems might replay from the beginning. None of these are defaults I accept. They are decisions made once, up front, because changing them later means coordinating every producer and consumer at once.

Partitioning for parallelism

Partition count and partition key choice are where most of the design work lives, because they determine ordering guarantees and how evenly load spreads across consumers at once. Kafka only guarantees ordering within a partition, so events that must be processed in order relative to each other need the same partition key, typically an entity ID. Get the key wrong and you either lose ordering you needed or, more commonly, create a hot partition where one key's volume dwarfs the rest and a single consumer becomes the bottleneck for the whole topic. I size consumer groups to partition count deliberately, since a group can never have more active consumers than partitions, and extra consumers just sit idle.

Idempotency or duplicates

At-least-once delivery is the practical default for Kafka, which makes duplicates a certainty over a long enough time window, whether from producer retries or consumer rebalances. I design consumers to be idempotent rather than trying to guarantee exactly-once delivery at the broker level: every handler produces the same end state whether an event is processed once or twice, usually via an idempotency key checked against a store before a side effect runs. That same property is what makes replay safe. If I ever need to reprocess a topic from an earlier offset to recover from a bug, idempotent consumers make that a non-event instead of a data-corruption risk.

Operating for throughput

Day to day, the two levers I watch are retention and consumer lag. Retention strategy (time-based, size-based, or compacted for topics that represent current state rather than a log of events) directly affects both storage cost and how far back a replay can reach. Consumer lag is the single most useful health signal a streaming system gives you: a consumer group falling behind its produced offsets is the earliest warning of a downstream bottleneck, well before anything looks broken from the outside.

Takeaways

  • Treat topic naming, schema evolution, and retention as explicit decisions, not defaults.
  • Partition key choice controls both ordering and load distribution. Get it wrong and you get hot partitions.
  • Design for at-least-once delivery: idempotent consumers make duplicates and replay safe.
  • Monitor consumer lag as your primary early-warning signal for the whole pipeline.