skip to content

A platform team wants to give internal consumers the low-callback benefits of event-carried state transfer for order events, but the order aggregate is large (megabytes, with attachments) and some fields are sensitive (PII, payment tokens) that shouldn't be broadcast to every subscriber. Design an approach that gets most of the coupling and freshness benefits of carried-state events without shipping the entire aggregate to everyone, and explain how it relates to event sourcing if the order service internally uses an event-sourced write model.

level: principalimportance: nice to knowfreq 30%

answer

  1. claim-check pattern: small payload + pointer to big/sensitive data
  2. size and sensitivity are separate problems, separate fixes
  3. internal event sourcing != external integration event shape
  4. anti-corruption-layer-style translation boundary
  5. transactional outbox as the stable public contract

basics

~20 s

Send a medium-sized event with the commonly needed fields plus a reference (like a claim check) that lets consumers who truly need the big or sensitive parts fetch them separately, rather than sending everything to everyone. Event sourcing, if used internally, is a separate decision about how the order service stores its own history - it doesn't have to leak into what gets published externally.

solid answer

~60 s

This calls for a hybrid, sometimes called the claim-check pattern: publish an event-carried state transfer payload containing the commonly needed, non-sensitive fields (status, totals, item summary) plus a reference/pointer (an ID or URL) to the full aggregate or to specific sensitive fields, stored separately (blob storage, a secured API). Most consumers get everything they need with zero callbacks; the rare consumer that needs the attachment or the payment token makes a targeted, authorized fetch only when it actually needs that data, keeping broadcast payloads small and access to sensitive fields auditable and gated. If the order service is internally event-sourced, that's an implementation detail of how it derives and persists its own current state from its private event log - it doesn't have to be the same event stream, format, or granularity as the external integration events it publishes; the service should translate its internal domain events into a deliberately designed, versioned public integration event (an anti-corruption-layer-style boundary), so internal refactors of the write model don't ripple out to every subscriber.

go deeper

for a junior

Not expected to design this; should recognize that sending everything to everyone can be wasteful or risky for sensitive data, at a conceptual level.

for a middle

Should be able to describe the basic idea of sending a small payload plus a reference for the rest, even without the 'claim check' name.

for a senior

Should propose the hybrid approach concretely, distinguish the size problem from the sensitivity problem, and note the residual coupling the reference/fetch reintroduces.

for a principal

Should additionally address the internal-event-sourcing-versus-external-contract boundary, name the anti-corruption-layer-style translation, and connect it to organizational-scale schema ownership and stability, such as via a transactional outbox.

## Two ways a pure carried-state design breaks down At the scale where a single aggregate is both large and partly sensitive, a pure event-carried state transfer design breaks down in two distinct ways, and the fix for each is different, which is exactly why this is a principal-level design problem rather than a simple binary choice between the two Fowler styles. - **The size problem** is that broadcasting megabyte-scale payloads (order attachments, generated PDFs, large line-item catalogs) to every subscriber on every change is wasteful: most consumers don't need the attachment at all, message brokers have practical size limits and cost implications for large messages, and network/storage costs scale with both payload size and subscriber count. - **The sensitivity problem** is separate: some fields (payment tokens, full PII) shouldn't be duplicated into every consumer's local storage at all, both because it multiplies the attack surface for a data breach and because it makes access auditing and revocation far harder - once a payment token is sitting in twelve different consumers' databases, you can no longer centrally control who can see it or expire it. ## The claim-check hybrid The standard resolution is a hybrid design often called the **claim-check pattern**, borrowed from the enterprise integration patterns literature: the event itself carries a right-sized payload of the fields most consumers commonly need (status, totals, a summarized item list) plus a reference - a claim check - pointing to where the full or sensitive data lives, such as an object-storage URL for a large attachment or an authorized API endpoint for payment details. This gets you most of the benefit of event-carried state transfer (no callback needed for the common case, which is the vast majority of consumer work) while avoiding both problems: - payloads stay small because bulk data isn't inlined, - and sensitive data stays behind an access-controlled boundary that can independently audit, rate-limit, and revoke access, rather than being copied wholesale into every subscriber's storage the moment an event fires. The rare consumer that genuinely needs the attachment or the token pays the cost of a targeted callback, but that's now the exception rather than the default path every consumer takes. ## Internal event sourcing versus the published contract The relationship to event sourcing is a separate axis that's easy to conflate with this but shouldn't be. If the order service's write model is internally event-sourced, that means its own current state is derived by replaying its private event log - this is an implementation choice about how the service persists and reconstructs its own aggregate, entirely internal to that service's boundary. Nothing about that internal choice requires the integration events published to other services to be the same events, at the same granularity, in the same format. In fact, coupling the internal event-sourcing log directly to external consumers is a well-known anti-pattern: - **Internal events** are usually fine-grained (`ItemAdded`, `DiscountApplied`, `AddressCorrected`) and reflect implementation details the service should be free to refactor, - whereas **a well-designed public integration event** is coarser-grained, deliberately versioned, and stable by contract because other teams depend on it. The right boundary is a translation layer - conceptually similar to an anti-corruption layer in Domain-Driven Design - where the order service listens to its own internal event-sourced stream and derives a smaller number of deliberately shaped, versioned public events (like the claim-check-style `OrderStateChanged` described above) to publish externally. This lets the internal write model evolve freely - adding new internal event types, changing internal aggregate boundaries, even switching away from event sourcing entirely - without any of that leaking into or breaking the public contract that dozens of other teams may depend on. ## Where it shows up A recognizable real-world analogue is how many systems separate their write-side domain events (used for internal replay, audit, and business logic) from a distinct 'integration events' or 'outbox' stream published externally via the **transactional outbox pattern** - the outbox table or topic is explicitly the translated, stable, versioned public contract, decoupled from whatever internal event shapes the write model happens to use. Getting this separation right is what lets a platform team scale an event-driven integration to dozens of consuming teams without every internal refactor becoming a cross-team breaking change, which is ultimately the organizational-scale version of the same coupling trade-off that shows up between a single producer and a single consumer at the micro level.

  • Why is coupling external consumers directly to an internal event-sourced log considered an anti-pattern?
    Internal event-sourced logs are typically fine-grained and reflect implementation details the owning team needs freedom to change, like splitting an event type or refactoring aggregate boundaries. If external consumers depend on that exact stream, every such internal refactor becomes a breaking change for other teams, which defeats much of the purpose of choosing event sourcing internally in the first place.
  • How would you decide which fields go in the claim-check event body versus behind the referenced fetch?
    Fields that most consumers commonly need and that are small and non-sensitive belong in the body, since inlining them is what eliminates the callback for the common case. Fields that are large, rarely needed, or sensitive belong behind the reference, since the goal is to make the exception path pay its own cost rather than making every consumer pay for the rare need.
  • What's the downside of the claim-check pattern compared to a pure event-carried state transfer design?
    Consumers that do need the referenced data still have a runtime dependency on that fetch succeeding, reintroducing some of the temporal coupling and freshness risk that carried-state was meant to avoid, just scoped to a smaller subset of consumers and fields. It's a deliberate trade-off, not a free win - you're choosing where to pay the callback cost rather than eliminating it entirely.

It's like a company sending a one-page executive summary to everyone (with a link to the full 200-page report in a secured archive for the few who need it), rather than mailing the entire report - and full folder of confidential attachments - to every single recipient.

saying these in an interview costs you the question

  • Proposes broadcasting the entire large/sensitive aggregate to every consumer to 'keep it simple'
  • Treats event sourcing and event-carried state transfer as the same thing
  • Assumes internal event-sourced events can be published externally unchanged with no translation layer
  • Doesn't separate the size problem from the sensitivity problem, applying one fix to both
  • Can't name a concrete drawback of the claim-check pattern itself

context