skip to content

A data catalog launched a year ago is now ignored because descriptions are stale and listed owners have left; how do you make it trusted again?

level: seniorimportance: should knowfreq 40%

answer

  1. fix ownership first
  2. automate everything a system knows
  3. focus on the most-used datasets
  4. show trust signals, hide the dead
  5. put it where people already work

basics

~20 s

Reassign ownership, automate every fact a system can supply, focus human effort on the most-used datasets, surface trust signals such as certification and freshness, deprecate dead assets, and embed the catalog in daily tools so it gets used and corrected.

solid answer

~50 s

I would treat it as a trust problem, not a tooling one. First **fix ownership**: map every dataset to a current team rather than a person, using org data and query logs, and make stale ownership a reported metric. Second, **automate** everything systems know — schemas, freshness, run history, usage, quality results, lineage — so those facts stay current as long as harvesting runs. Third, **focus** human documentation on the datasets that carry most of the query volume, using usage data to pick them. Fourth, add **trust signals**: certified versus unreviewed, last refresh, check status, top users — and **deprecate or hide** unused and broken assets so search results are clean. Finally, **embed** the catalog where people already work — links from dashboards and query editors, required entries when publishing a dataset — so it is used, and usage produces corrections.

go deeper

for a junior

Know why a catalog loses trust when descriptions and owners go stale.

for a middle

Explain which metadata can be automated and why focusing on the most-used datasets matters.

for a senior

Lay out a recovery plan with team ownership, trust signals, deprecation and workflow integration, and the metrics to track it.

for a principal

Decide the organisational mechanisms, such as publishing requirements and ownership reporting, that stop the catalog decaying again.

## Diagnosing why catalogs die Catalogs usually fail the same way: a big launch, a push to document everything, then decay. Descriptions go stale, owners move on, search returns hundreds of near-duplicate tables, and users learn that asking a colleague is faster. The fix is to make the catalog **accurate by default** and **useful in daily work**, not to relaunch it. ## A recovery plan 1. **Ownership to teams, not people.** Assign each dataset to an owning team, derived from who builds it (pipeline code owners) and who maintains it. Individuals change roles; teams persist. Publish the share of datasets whose owner is unknown and drive it down. 2. **Automate the facts.** Everything a system can supply should be harvested continuously: schemas, refresh times, run history, row counts, quality-check results, query counts, lineage. They stay current as long as harvesting runs, and they are what users check first. 3. **Prioritise by usage.** Query logs show that a small share of datasets carries most of the reads. Document those well; let the long tail keep automated metadata and suggestions. 4. **Add trust signals.** Show a **certification** state (reviewed by the owner as the recommended source), freshness, quality status, and popularity. Users decide quickly when they can see these. 5. **Clean the search.** Mark unused, broken or superseded datasets as deprecated and demote or hide them, so search results point at what should be used. 6. **Embed in workflows.** Link from dashboards and query editors to the dataset's catalog page; require a catalog entry (owner, description, classification) before a dataset is published to shared spaces. 7. **Close the feedback loop.** Let users flag wrong information and route it to the owner; measure time to fix. ## Measuring recovery | Metric | What it tells you | |---|---| | Datasets with a valid owning team | whether accountability exists | | Top datasets with a reviewed description | whether human effort went where usage is | | Catalog searches and page views per week | whether people use it | | Share of reads hitting certified datasets | whether trust signals steer usage | | Open user-reported issues and time to fix | whether the loop works | ## What not to do - Another documentation sprint across every table. - A mandate to use the catalog without making it better than asking colleagues. - Keeping every dataset visible in search because deleting feels risky. ## Why interviewers ask it Many candidates can list catalog features; fewer have seen one fail. A senior answer names the **root causes** (ownership decay, manual metadata, noise), prioritises **automation and usage data**, and measures adoption rather than coverage.

  • Why assign ownership to teams rather than individuals?
    Individuals change roles and leave, which is exactly how owner fields went stale. A team persists through staff changes, and team membership is already maintained elsewhere, so the catalog can resolve the current contacts from it.
  • How do you decide which datasets to certify first?
    By usage and business importance — the datasets behind executive and regulatory reporting, and those with the most distinct readers. Certification requires the owner to review definitions and checks, so it is spent where it steers the most decisions.

saying these in an interview costs you the question

  • Relaunching the catalog with another manual documentation push
  • Listing a person who left as the owner indefinitely
  • Keeping deprecated and broken datasets equally visible in search
  • Measuring success by the number of documented tables instead of usage