skip to content

How should you design Kafka topics for domain events — granularity (one event type per topic vs many), keying, and retention/compaction — and what is a domain event versus an integration event?

level: middleimportance: should knowfreq 55%

answer

  1. Domain event = internal fact; integration event = curated public contract
  2. Key by aggregate id → same partition → per-entity order
  3. One-type-per-topic = clean but no cross-type order + sprawl
  4. Related types on one topic = ordered lifecycle (RecordNameStrategy)
  5. compact for state/ECST, delete for notifications
  6. Partitions = parallelism; resizing reshuffles keys

basics

~20 s

A domain event captures something meaningful in the business domain (OrderPlaced). Design topics around an aggregate/entity, key events by the aggregate id for ordering, choose one-type-per-topic for clarity or grouped types for related events, and use compaction for state-style events and time retention for pure notifications.

solid answer

~50 s

A **domain event** represents a business-meaningful change within a bounded context (e.g. `OrderPlaced`); an **integration event** is the version you publish *across* context/service boundaries — often a curated, stable, public-contract subset of the internal domain event. For topic design: scope topics around an **aggregate/entity type** (e.g. an `orders` event topic), and **key by the aggregate id** so all events for one entity share a partition and stay ordered. Granularity is a trade-off: **one event type per topic** gives clean schemas, independent retention, and easy subscription, but proliferates topics and loses cross-type ordering; **multiple related types on one topic** (with `TopicRecordNameStrategy` in the registry) preserves ordering across those types and reduces topic sprawl but couples them. For retention: use **time/size retention** (`cleanup.policy=delete`) for transient notifications, and **log compaction** (`cleanup.policy=compact`) for state-snapshot/ECST events so consumers can rebuild current state by replay. Keep partitions sized for throughput and consumer parallelism, and version schemas via a registry.

go deeper

for a junior

Know that domain events are business facts and you key topics by the entity id.

for a middle

Weigh one-type-per-topic vs grouped, and pick compaction vs delete retention appropriately.

for a senior

Reason about integration vs domain events, subject-naming strategies, and partition-count consequences.

for a principal

Define org topic taxonomy, public integration-event contracts, and retention/partition standards across teams.

## Domain event vs integration event - A **domain event** is a fact that matters **inside a bounded context** — it's part of your domain model: `OrderPlaced`, `PaymentCaptured`. It may be rich and reflect internal structure. - An **integration event** is what you publish **across boundaries** to other services/teams. Best practice is to **not** leak your raw internal domain event; instead publish a **deliberately designed, stable, public** integration event (a curated subset/translation) so internal refactors don't break external consumers. The integration event is a **published API**; the domain event is internal. ## Topic granularity — the central design choice **Option A: one event type per topic** (`order-placed`, `order-cancelled`). - Pros: clean single schema per topic; per-type retention/compaction; consumers subscribe only to what they need; simplest schema-registry story (TopicNameStrategy). - Cons: **no ordering across types** (a cancel could be processed before its place if on separate topics/partitions); topic **proliferation**. **Option B: many related event types on one topic** (an `orders` topic carrying placed/updated/cancelled). - Pros: **ordering preserved** across all events for an entity (same topic, keyed by entity → same partition); fewer topics; a consumer sees the full lifecycle in order. - Cons: heterogeneous schemas on one topic (needs `RecordNameStrategy`/`TopicRecordNameStrategy` so each type has its own schema subject); consumers must filter types they don't care about; retention policy is shared. **Guideline**: group event types that belong to the **same aggregate and need mutual ordering** onto one topic; split unrelated streams. Don't make 'one giant topic for everything' (loses parallelism, forces global filtering) nor 'a topic per field change' (sprawl). ## Keying **Key by the aggregate/entity id** (orderId, customerId). Kafka's default partitioner hashes the key, so all events for one entity land on the **same partition**, and Kafka guarantees **order within a partition** — giving per-entity ordering, which is what business correctness usually needs. Null keys → round-robin → no ordering. Keying also makes **log compaction** meaningful (latest value per key). ## Retention & compaction - `cleanup.policy=delete` + `retention.ms` / `retention.bytes`: time/size-based retention for **transient notifications** (you only care about recent events). - `cleanup.policy=compact`: **log compaction** keeps the **latest** record per key indefinitely — ideal for **state/ECST** events so a consumer can replay the topic and **materialize current state** of every entity (foundation of `KTable`). A **tombstone** (null value) deletes a key. - You can combine (`compact,delete`) for compacted topics that still age out tombstones. ## Partition count Partitions set the **max consumer parallelism** in a group (one partition per consumer at most) and affect ordering granularity. Over-partitioning wastes resources and increases end-to-end latency/rebalance cost; under-partitioning caps throughput. Choose for projected throughput and consumer count, and remember **increasing partitions later changes key→partition mapping** (breaks existing ordering and compaction grouping), so size with foresight. ## Pitfalls - Publishing raw internal domain events as the cross-team contract (brittle). - Splitting lifecycle events across topics when they need mutual ordering. - Null/poorly-chosen keys destroying ordering and breaking compaction. - Forgetting that adding partitions later re-shuffles key placement.

  • Why might you put OrderPlaced, OrderUpdated, and OrderCancelled on one topic rather than three?
    To preserve ordering across the whole order lifecycle: with one topic keyed by orderId, all three event types for an order land on the same partition and are processed in order. Three separate topics give no cross-type ordering guarantee. The cost is heterogeneous schemas (use TopicRecordNameStrategy) and consumer-side filtering.
  • Why not just publish your internal domain event directly to other teams?
    It couples external consumers to your internal model, so refactors break them. Publish a curated, stable integration event (a public-contract subset/translation) instead, decoupling internal evolution from the external API.

saying these in an interview costs you the question

  • Using null/random keys when per-entity ordering is needed
  • Treating internal domain events as the cross-team public contract
  • Assuming you can freely increase partition count without affecting ordering/compaction
  • Using delete retention for state-snapshot events that consumers need to replay
  • One mega-topic for all events, forcing global filtering and killing parallelism

context