In a data mesh, what distinguishes a data product from a table a team happens to publish?
answer
- a named owner accountable for it
- consumers treated as customers
- discoverable, addressable, self-describing
- trustworthy, with stated SLOs
- interoperable and secure by standard
basics
~20 sA data product has an accountable owner and is built for consumers - discoverable, addressable, self-describing, trustworthy with stated objectives, interoperable through global standards and secured by global access rules. A published table has none of those guarantees.
solid answer
~40 sA table someone publishes is just data at a location. A **data product** is data served **for consumers** with someone accountable for it. In the 2019 article that introduced data mesh, a domain data product has baseline qualities: **discoverable** (registered in a catalog with owner, origin, lineage and samples), **addressable** (a stable, conventional address for programmatic access), **trustworthy and truthful** (the owner commits to service-level objectives on its quality and timeliness), **self-describing** (clear semantics and syntax, ideally with examples), **interoperable** under global standards, and **secure** under global access control. A **data product owner** measures consumer satisfaction and lead time to use, and the product includes its code, data, metadata and infrastructure as one deployable unit. Without those, it is a dump.
go deeper
Know that a data product has an owner and is built for its consumers, unlike a table that merely exists.
Explain the baseline qualities such as discoverable, addressable, trustworthy and self-describing, with an example of each.
Show how ownership, objectives and measures such as lead time to use are set up for a real product.
Decide which datasets the organisation invests in as products and how product ownership is funded and staffed in domains.
## From dataset to product In a centralised platform, a table appears because a pipeline wrote it. Nobody promises it will keep its shape, refresh on time or mean what its name suggests. Data mesh's second principle, **data as a product**, asks domains to treat their analytical data as something **consumers depend on** — with a named owner and explicit qualities. ## The baseline qualities The 2019 article lists qualities a domain data product needs to delight its consumers: | Quality | What it requires in practice | |---|---| | **Discoverable** | registered in a central catalog with owner, source of origin, lineage and sample data | | **Addressable** | a unique, stable address following a global naming convention | | **Trustworthy and truthful** | the owner states service-level objectives for accuracy and timeliness and tests them at creation | | **Self-describing semantics and syntax** | schema and meaning documented, ideally with sample data, so consumers need no hand-holding | | **Interoperable, governed by global standards** | shared identifiers and formats so products can be joined | | **Secure, governed by global access control** | access policies applied consistently across products | The 2020 article refers back to that list as capabilities including discoverability, security, explorability, understandability and trustworthiness. ## Ownership and roles - A **domain data product owner** is accountable for the product: who its consumers are, how they use it, and objective measures such as data quality, lead time to use and user satisfaction. - **Data product developers** in the domain build, maintain and serve it. - The product is the **architectural quantum**: its code, its data and metadata, and its infrastructure, deployable independently. ## A test to apply Ask of any dataset: 1. Who is accountable if it is wrong, late or changes shape? 2. Can a stranger find it and understand it without asking anyone? 3. What objectives does it commit to, and are they measured? 4. Can it be joined with other domains' products on shared identifiers? A published table usually fails all four; a data product answers them. ## Boundaries with neighbouring practices - The **contract** that formalises a product's schema and objectives, and how it is enforced in CI, is data-contract practice. - **Catalog** mechanics — how the product is registered and found — belong to catalog practice. - Data as a product is the **ownership and intent** behind both. ## Why interviewers ask it Many organisations use "data product" loosely for any table. Interviewers check whether a candidate can say what **the term demands**: accountable ownership and consumer-facing qualities, with objectives the owner is measured on.
- Who should measure whether a data product is successful, and by what?Its data product owner, using measures the consumers feel — data quality against stated objectives, lead time for a new consumer to start using it, and consumer satisfaction — rather than only pipeline uptime.
- Why does addressability need a global convention rather than each domain's choice?Consumers access many products programmatically. If every domain names and locates its data differently, each consumer must learn each convention, which defeats self-service; a shared convention makes any product reachable the same way.
saying these in an interview costs you the question
- Calling any published table a data product
- Leaving data product ownership with a central team by default
- Promising quality without stated and measured objectives
- Letting each domain invent its own identifiers for shared entities