Compare fat (event-carried state transfer) events versus thin (notification) events. What are the trade-offs, and when would you choose each for a Kafka topic?
answer
- Thin = id only, call back; Fat = full state embedded
- Fat = ECST = autonomous consumers, local read model
- Fat pairs with log compaction / KTable
- Fat costs: size, PII duplication, staleness
- Claim-check = thin + pointer for big payloads
basics
~20 sA thin event carries only an id/reference, so consumers must call back to fetch details. A fat event carries the full state, so consumers need no callback but the payload is larger and may go stale. Fat reduces coupling/load; thin keeps payloads small and authoritative.
solid answer
~60 sA **thin (notification) event** carries minimal data — typically just an entity id and event type (e.g. `OrderUpdated{orderId}`). Consumers must call back to the source service or DB to fetch current state. This keeps payloads tiny and always points to the authoritative source, but creates a query fan-out and runtime coupling back to the producer. A **fat event (event-carried state transfer, ECST)** embeds the full relevant state in the payload (`OrderUpdated{orderId, items, total, status, ...}`). Consumers can build and serve their own local read models without calling back, which removes synchronous coupling and lets them survive producer downtime. Costs: larger messages, potential PII duplication, and snapshots that are stale relative to the source. In Kafka, ECST pairs naturally with **log compaction** (keyed topic retains the latest state per key) so a consumer can rebuild a full materialized view by replaying the topic. Choose fat when you want autonomous consumers and decoupling; thin when payloads are huge, data is sensitive, or freshness from a single source of truth is critical.
go deeper
Know that thin = id + callback, fat = full state in the payload.
Articulate the coupling/staleness/size trade-offs and pick one per scenario.
Tie ECST to compaction, versioning for out-of-order snapshots, and claim-check for large payloads.
Define org policy on payload shape, PII-in-events governance, and compaction/retention strategy across topic families.
## The two shapes of an event payload When you publish an event that 'something changed', you must decide **how much data** to put in it. ### Thin event (a.k.a. notification event) Carries only enough to identify *what* changed: an id, maybe the event type and timestamp. Example: `CustomerUpdated{customerId: 42}`. - **Consumer flow**: receive event → call the source service's API (or query a shared DB) → get current details. - **Pros**: tiny messages; the data fetched is always **current**; no duplication of sensitive fields on the wire/log. - **Cons**: every consumer that needs detail must **call back**, creating a query storm (N consumers × M events) and **runtime coupling** — if the producer is down, consumers are blocked. You also lose the ability to replay history meaningfully, because the callback returns *today's* state, not the state at event time. ### Fat event (event-carried state transfer, ECST) Embeds the full relevant state: `CustomerUpdated{customerId: 42, name, tier, address, ...}`. - **Consumer flow**: receive event → update a **local copy/read model** → serve queries from that copy. - **Pros**: consumers become **autonomous** — no callback, they survive producer outages, and they can keep their own denormalized view tuned for their queries. Strong decoupling. - **Cons**: bigger payloads; **data duplication** (including possibly PII — a governance concern, e.g. GDPR right-to-erasure now spans many local copies); the embedded snapshot is **stale** the instant a newer event exists; and you must handle **out-of-order/old** snapshots (carry a version/sequence so consumers ignore older state). ## Kafka-specific mechanics - **Log compaction** (`cleanup.policy=compact`) is the killer feature for ECST: with a topic keyed by entity id, Kafka retains at least the **latest** value per key forever. A new consumer can replay the compacted topic from the beginning and **materialize the full current state** of every entity — the topic *is* the dataset. This is the foundation of Kafka Streams `KTable` and the 'turning the database inside out' idea. - **Tombstones**: a record with a null value on a compacted topic deletes that key — important for ECST erasure. - **Message size**: very large fat events can hit `max.message.bytes` / `max.request.size`; for huge blobs use the **claim-check** pattern (event carries a pointer to object storage), which is a deliberate middle ground. ## Hybrid / middle grounds - **Claim-check**: thin event + pointer to a large payload in S3/blob store. - **Delta/partial-fat**: carry the changed fields plus key, not the whole entity. ## Decision guide Choose **fat/ECST** when: you want autonomous consumers, resilience to producer downtime, replayable materialized views, and the entity isn't enormous. Choose **thin** when: payloads would be huge, data is highly sensitive and you want one authoritative copy, or consumers genuinely need the freshest possible value at read time. ## Common mistakes - Treating staleness in fat events as a bug rather than an inherent, version-managed property. - Forgetting to include a **version/sequence** so consumers can discard older snapshots that arrive out of order. - Putting PII in fat events without a deletion/compaction-tombstone story.
- Which Kafka feature makes event-carried state transfer especially powerful, and why?Log compaction (cleanup.policy=compact) on a keyed topic retains the latest value per key, so a new consumer can replay the topic and rebuild a full materialized view of current state. This underpins Kafka Streams KTables.
- Fat events duplicate data into many consumers. What governance problem does that create and how do you handle it?Sensitive data (PII) is now copied across consumer stores, complicating deletion/GDPR erasure. Handle it with compaction tombstones (null-value records) to remove keys, plus consumer-side deletion of derived copies, or keep such fields thin.
saying these in an interview costs you the question
- Claiming fat events are always better with no downsides
- Ignoring that fat-event snapshots can be stale and need versioning
- Forgetting log compaction as the enabler for ECST materialized views
- Not recognizing PII duplication / erasure as an ECST governance concern