In a data mesh, how do source-aligned and consumer-aligned data products differ, and who should own a dataset that combines several domains?
answer
- facts as the business records them
- reshaped for a group of uses
- aggregates change more often
- owner is whoever serves its consumers
- regenerate from the sources
basics
~20 sSource-aligned products publish a domain's business facts close to how they happen; consumer-aligned products reshape and combine them for a group of use cases. A cross-domain dataset should be owned by a team whose purpose is serving that use, not by one source domain.
solid answer
~50 s**Source-aligned** data products expose the facts a domain's operational systems record — orders placed, streams played — cleaned and deduplicated by that domain, and they change slowly because they mirror the business. **Consumer-aligned** (or aggregate) products **transform and combine** source products into structures fitting a closely related group of uses — a customer view for marketing, a social graph for recommendations — and they change more often as use cases evolve. For a dataset combining several domains, ownership should sit with a team **whose purpose is to serve that use**: either the consuming domain itself, or a newly formed team when several consumers share it. Assigning it to one source domain overloads that team with concerns it does not own. Consumer-aligned products should be **regenerable from the source products**, so they can be rebuilt as needs change.
go deeper
Know that some data products publish business facts and others reshape them for particular uses.
Explain how source-aligned and consumer-aligned products differ in content, change rate and ownership.
Decide ownership for a cross-domain dataset and design it to be regenerated from published source products.
Set the organisation's rules for when aggregates get their own team and how orphaned products are prevented.
## Two kinds of domain data The 2019 article that introduced data mesh distinguishes: - **Source-oriented (source-aligned) domain data**: the analytical form of facts the business records in its operational systems. A source domain cleans, deduplicates and enriches its own events so other domains can consume them **without repeating that cleansing**. These products reflect reality and change slowly. - **Consumer-oriented and shared domain data**: data shaped for a **closely related group of use cases**. The article's example is a recommendations domain building a graph of users' social connections, which a notifications domain might also use — at which point a "user social network" can become a shared domain dataset of its own, curated by its own team. ## How they differ | Aspect | Source-aligned | Consumer-aligned (aggregate) | |---|---|---| | Content | business facts as recorded | transformed, combined, aggregated views | | Rate of change | slow; mirrors the business | frequent; follows use cases | | Owner | the source domain | the consuming domain or a dedicated team | | Rebuild | from operational systems | from source-aligned products | | Example | orders placed, payments captured | customer 360, churn features, social graph | ## Who owns a cross-domain dataset A "customer 360" joins orders, support tickets and marketing interactions. Candidates for ownership: 1. **One source domain** (for example orders) — usually wrong: it becomes responsible for other domains' semantics and for consumers it does not serve. 2. **The main consuming domain** (for example marketing) — right when one domain is the dominant user and understands the need. 3. **A new dedicated team** — right when several domains depend on it and it has become a product in its own right. The rule of thumb: ownership follows **who serves the consumers and can be held to the product's objectives**, and the product depends on **source-aligned products**, not on raw operational data from other domains. ## Regenerability Because consumer-aligned products change with use cases, the article notes a platform should make it **easy to regenerate** them from source products. Keeping them derived — not hand-maintained copies — lets them be rebuilt when definitions change and keeps the source domains the single place facts are corrected. ## Failure modes - A consumer-aligned product reads another domain's **internal** tables instead of its published product, coupling to internals. - Aggregates become **orphans** after the project that built them ends; nobody owns their objectives. - Source domains are pressured to add consumer-specific shaping into source products, which bloats them. ## Why interviewers ask it Ownership of combined datasets is where decentralised ownership gets difficult in practice. A senior answer distinguishes the two kinds, assigns ownership by **who serves the use**, and insists on building from **published source products**.
- Why should a consumer-aligned product read source domains' published products rather than their internal tables?Published products carry the owner's commitments on shape, meaning and quality. Internal tables can change without notice, so depending on them couples the aggregate to another domain's implementation and makes it break unpredictably.
- When would you create a new team to own an aggregate data product?When several domains rely on it, its logic is substantial, and no single consumer is the natural owner. At that point it behaves like a product of its own and needs a team accountable for its objectives.
saying these in an interview costs you the question
- Assigning every cross-domain dataset to the domain that supplies most of its rows
- Building aggregates on other domains' internal tables
- Pushing consumer-specific shaping into source-aligned products
- Leaving aggregate products without an owner once the project ends