What minimum metadata should a data platform require before a dataset can be published for shared use, and how do you enforce it without stalling teams?
answer
- a short list: can I use it?
- owner, description, classification, freshness
- enforce at publish, in the pipeline
- tiers for shared versus personal
- measure friction and fix the path
basics
~20 sRequire only what a consumer needs to decide whether to use a dataset — owning team, description, classification, update expectation — enforce it automatically at publication, apply it only to shared tiers, and make compliance the easiest path.
solid answer
~40 sI would set a **short** mandatory list aimed at the consumer's question, *can I use this?*: an **owning team**, a **description** that says what one row represents, a **sensitivity classification**, and an **update expectation** (how often it refreshes). Anything more — column descriptions, glossary links, quality checks — becomes a requirement for **certified** status rather than for publication. Enforcement should be **automatic and early**: a check in the pipeline or publishing step that fails when required fields are missing, instead of a review queue. Apply it only to **shared** spaces; personal and sandbox areas stay free, so experiments are not taxed. And make compliance cheap: templates, pre-filled suggestions, owners defaulted from the repository. I would measure blocked publications and time to comply, and loosen or automate whatever causes most friction.
go deeper
Know that a platform can require basic metadata, such as an owner and a description, before data is shared.
Explain which fields help a consumer decide whether to use a dataset, and how a pipeline check can enforce them.
Design tiers and automated defaults that keep the standard enforceable without manual review.
Own the trade-off between catalog quality and adoption, choose what is mandatory at each tier, and measure the friction the standard creates.
## The judgment call Every catalog programme hits the same tension. **Too little** required metadata and the catalog fills with undocumented, ownerless tables nobody can trust. **Too much**, enforced by reviewers, and teams route around the platform — publishing into personal spaces, exporting to spreadsheets — which is worse than an incomplete catalog. There is no single right list; the decision is where to put the line and how to enforce it. ## A tiered standard | Tier | Who can create | Required metadata | Enforcement | |---|---|---|---| | **Sandbox / personal** | anyone | none beyond automatic technical metadata | none; not discoverable to others by default | | **Shared** | teams | owning team, dataset description (what one row is), sensitivity classification, update expectation | automated check at publish time | | **Certified** | owners, after review | column descriptions for key columns, glossary links, quality checks, documented lineage | owner sign-off, periodic re-review | The shared tier's list is deliberately short: each field answers part of **"can I use this, and whom do I ask?"**. ## Enforcement that does not stall teams 1. **Shift left.** Declare metadata alongside the pipeline code (a small file next to the model), and validate it in the same CI check that builds the pipeline. 2. **Automate defaults.** Owning team from the repository's code owners; classification suggested by scans; refresh expectation from the schedule. 3. **Fail with a fix.** A failed check should say exactly which field is missing and offer the template. 4. **Grace for migration.** Existing datasets get a deadline and a dashboard of gaps, not an instant block. 5. **Measure the friction.** Track publications blocked, time to resolve, and datasets appearing in unmanaged places; if the last number rises, the standard is too heavy. ## What to argue about - **Classification at publication or later?** Requiring it up front prevents unprotected sensitive data but adds work; a common compromise is to default to the restrictive class until reviewed. - **Quality checks for shared data?** Mandatory checks raise trust but slow publishing; many platforms require them only for certified datasets. - **Who can grant exceptions?** A named platform owner with a time limit, recorded in the catalog. ## Why interviewers ask it It is a principal-level trade-off with no single answer: the candidate must balance catalog quality against adoption, pick enforcement mechanisms, and show how they would **measure** whether the standard helps or hurts.
- Why not require column-level descriptions for every shared dataset?Because the cost falls on every publication while most columns are rarely read. Requiring them for certified datasets puts the effort where usage and business importance justify it, and suggestions can cover the rest.
- What signal tells you the standard is too heavy?Growth in data published outside the managed spaces — personal areas used for shared work, exports, unmanaged copies — alongside rising time to publish. Teams are voting with their feet, and the standard should be simplified or automated further.
saying these in an interview costs you the question
- Requiring exhaustive documentation before any dataset can be shared
- Enforcing metadata through a manual review queue
- Applying the same rules to personal sandboxes as to certified data
- Measuring only compliance, never the friction it causes