What are the four principles of data mesh, and what problem of centralised data platforms is each one meant to solve?
answer
- who owns it
- how it is served
- what makes it easy
- how standards hold
basics
~20 sDomain ownership moves analytical data to the teams that know it; data as a product makes it usable by others; a self-serve platform makes that affordable for every domain; federated computational governance keeps it interoperable through shared rules the platform enforces.
solid answer
~50 sData mesh, as described by Zhamak Dehghani, rests on four principles. **Domain-oriented ownership**: analytical data is owned by the business domains that produce and understand it, instead of a central data team that becomes a bottleneck and lacks domain knowledge. **Data as a product**: each domain serves its data as a product with consumers treated as customers — discoverable, understandable, trustworthy — to avoid silos and unusable dumps. **Self-serve data platform**: a platform team provides the infrastructure so that domains can build and run data products without each becoming infrastructure experts. **Federated computational governance**: a federation of domain and platform owners agrees on global rules — identifiers, interoperability, security — and the platform enforces them automatically, so decentralisation does not become chaos. The principles are meant to work **together**; adopting only the first produces distributed silos.
go deeper
Be ready to name the four principles and say in one sentence what each changes.
Explain the problem each principle addresses and why adopting only some of them fails.
Describe how the central data team's role changes and what governance as platform-enforced rules looks like.
Judge whether the organisation's scale and domain structure justify the approach, and plan the order of adoption.
## The problem data mesh responds to In a classic centralised architecture, operational teams emit data, a **central data team** ingests it into a lake or warehouse, cleans it and serves it to analysts. As sources and consumers multiply, that team becomes a **bottleneck**: it lacks the domain knowledge to model data correctly, every change queues behind it, and quality problems are fixed far from where they arise. Data mesh, introduced by Zhamak Dehghani in articles published on martinfowler.com in 2019 and 2020, proposes a decentralised alternative built on **four principles**. ## The four principles | Principle | What it means | Problem it targets | |---|---|---| | **Domain-oriented decentralised data ownership** | the domains that generate or understand data own its analytical form | central team bottleneck and missing domain knowledge | | **Data as a product** | each domain serves data as a product with an owner, for consumers treated as customers | unusable, untrusted, undiscoverable data dumps and silos | | **Self-serve data infrastructure as a platform** | a platform team offers the tooling to build, run and discover data products | cost and duplicated effort of every domain building infrastructure | | **Federated computational governance** | a federation of domain and platform owners sets global rules; the platform enforces them automatically | inconsistency and non-interoperable data once ownership is decentralised | ## Why they depend on each other The 2020 article states the principles are intended to be **collectively necessary**: - Decentralised ownership **without** data-as-a-product produces many silos instead of one. - Data as a product **without** a self-serve platform makes each domain pay the full infrastructure cost, which few will. - All of the above **without** federated governance produces data that cannot be joined across domains — different customer identifiers, incompatible formats. ## What changes in practice 1. **Accountability for quality moves upstream**, as close to the source as possible. 2. **Pipelines become internal to domains** rather than a central shared layer. 3. **The central team becomes a platform team**, building capabilities rather than datasets. 4. **Governance becomes code**: policies such as access control and classification are implemented by the platform, not enforced by review meetings. ## Common misunderstandings - Data mesh is **not a product** or a technology; it is an organisational and architectural approach. - It does **not remove central work**; it moves it into the platform and the governance federation. - It is **not required** for every organisation — the approach targets scale and complexity that small organisations may not have. ## Why interviewers ask it Data mesh is a common talking point in data engineering and analytics interviews. A good junior answer names the four principles precisely and ties each to the problem it solves; better answers add that they only work **together**.
- What happens if an organisation adopts domain ownership but skips the self-serve platform?Each domain must build and run its own storage, pipelines, access control and publishing. Most lack the skills or budget, so data products are built inconsistently or not at all, and the central bottleneck is replaced by many small ones.
- Does data mesh mean there is no central data team?No. The central team usually becomes the platform team that builds self-serve capabilities, and central roles take part in the governance federation. What moves to domains is ownership of the data itself.
saying these in an interview costs you the question
- Describing data mesh as a technology or a product to install
- Naming only decentralised ownership and ignoring the other three principles
- Claiming data mesh removes the need for any central platform team
- Treating federated governance as no governance