skip to content

A team praises their choreographed order-fulfillment flow as 'loosely coupled' because services only communicate via published events with no direct calls between them. Six months later, changing the shape of the OrderPlaced event breaks three downstream services nobody remembered were listening. What went wrong with the 'loosely coupled' claim?

level: middleimportance: must knowfreq 70%

answer

  1. coupling moves, doesn't vanish
  2. schema is the new API
  3. no registry of subscribers
  4. schema registry + compatibility checks
  5. consumer-driven contract tests

basics

~20 s

Choreography removes direct service-to-service calls, but every subscriber still depends on the exact shape of the events it listens to. That's still coupling - to a shared event contract instead of an API - and it's invisible because nobody 'calls' anywhere, so nobody tracks who depends on what.

solid answer

~30 s

Choreography eliminates temporal/call coupling - no service blocks waiting on another's API - but it does not eliminate contract coupling: every consumer is coupled to the event's schema and its implied semantics. Because there's no central registry of 'who listens to what' the way an orchestrator's explicit process definition provides, this coupling is invisible until you try to change the event and something downstream breaks. It's sometimes called implicit or event-schema coupling, and it's worsened when consumers infer business meaning from fields the producer never intended as a stable contract.

go deeper

for a junior

Can recognize that removing direct calls doesn't mean removing all dependency.

for a middle

Names the concept as schema/contract coupling and explains why it's harder to discover than API coupling.

for a senior

Proposes concrete mitigations - schema registry with compatibility rules, consumer-driven contracts, additive-only evolution - and explains the production symptom (delayed, data-quality-style incidents).

for a principal

Treats event schemas as a cross-team API-governance problem, with versioning policy and deprecation process owned organizationally, not just technically.

## Coupling changes shape, it does not disappear Coupling doesn't disappear in choreography — it changes shape. The team's mistake was equating "no direct method call" with "no dependency." In reality, every service that subscribes to `OrderPlaced` has, by definition, taken a dependency on that event's schema: - its field names, - types, - required-vs-optional fields, - and any implicit meaning baked into specific values. When the producer renames a field, changes a type, or removes something it thought was unused, every consumer that referenced it breaks at runtime — often with a deserialization error or, worse, silent misbehavior if the field was optional and the consumer just used a stale default. ## Why this is worse than an ordinary API dependency The reason this coupling is so much more dangerous than an ordinary API dependency is **discoverability**. - If service B calls service A's REST endpoint directly, A's team can usually find B in a service registry, an API gateway's access logs, or a dependency graph, and warn B before making a breaking change. - With events, the producer publishes to a topic and has no built-in way to know who's subscribed — subscriptions are configured entirely on the consumer side, often in config files or code that the producer's team never sees. The result is coupling that is real but untracked: it only becomes visible reactively, when a change ships and pages start firing in unrelated services. ## Fan-out compounds it This problem compounds with fan-out. A popular event like `OrderPlaced` tends to accumulate consumers over time — fraud detection, loyalty points, analytics, email receipts, a new team's recommendation engine — each added independently, each adding a hidden dependency edge that isn't visible from the producer's code. The producer's small, well-intentioned refactor (e.g., splitting a `shippingAddress` string into structured fields) can silently break five services at once, and diagnosing the breakage means separately tracing failures back through logs in each of those five services rather than a single failing build. ## Making the implicit contract explicit The standard mitigations all aim at making the implicit contract explicit and stable. 1. **Treating the event schema as a first-class, versioned public contract** — with a schema registry (e.g. enforcing Avro/Protobuf/JSON-Schema compatibility rules), additive-only changes, and deprecation windows before removing fields — turns a silent breaking change into a compile-time or registration-time failure instead of a production incident. 2. **Consumer-driven contract testing** (each consumer publishes the subset of the schema it actually relies on, and CI runs those contract tests against the producer's build) restores some of the discoverability an API-based dependency would give for free. 3. Some teams also deliberately **keep "wide" events lean** — publishing a stable minimal fact plus a reference ID, and letting interested consumers fetch full detail via a query API — to shrink the surface area that can break. ## The failure mode in production The failure mode in production is rarely a crash at the producer; it's a slow-motion incident at the consumers, discovered hours or days later as data-quality drift ("why did loyalty points stop accruing for orders since Tuesday?") rather than an immediate alert, precisely because the producer's own tests and deploy pipeline had no idea the consumer existed. A well-known real-world pattern that grew directly out of this pain is the widespread adoption of schema registries alongside Kafka — Confluent's Schema Registry being the most common example — specifically to enforce backward/forward compatibility checks on topics before a producer is allowed to publish an incompatible schema, turning what used to be an invisible six-months-later break into a rejected deployment.

  • What's a concrete way to make a breaking event-schema change safely in a choreographed system?
    Publish the new field or shape alongside the old one (additive change) for a deprecation window, let consumers migrate on their own schedule by watching for the new fields, then remove the old fields only after confirming via the schema registry or usage metrics that nothing still depends on them. This mirrors how you'd version a public REST API rather than mutating it in place.
  • Does adding an orchestrator eliminate this schema-coupling problem?
    Not entirely - the orchestrator's command messages to participants and any events it emits are still contracts consumers depend on. What orchestration adds is a single, known list of participants the orchestrator calls directly, which at least makes the dependency graph explicit and discoverable, unlike choreography's hidden subscriber list.

It's like a radio station broadcasting on a frequency: the station doesn't know how many radios are tuned in, so it feels free of obligation - but change the broadcast format and every listener's radio stops working at once, with no way for the station to have warned them first.

saying these in an interview costs you the question

  • Claims choreography has 'no coupling' at all
  • Doesn't recognize the event schema itself as a contract
  • Assumes a producer can safely change a field just because no direct callers exist
  • Has no answer for how to find out who consumes a given event before changing it

context