In dbt, what do dbt docs generate and dbt docs serve produce?
answer
- one command writes files, the other shows them
- two JSON artifacts, one needs the warehouse
- long descriptions live in Markdown blocks
- manifest.json, catalog.json, doc()
basics
~20 sdbt docs generate writes manifest.json and catalog.json into the target directory — the project graph plus column metadata queried from the warehouse. dbt docs serve starts a local web server rendering those files as a browsable catalog with a lineage graph.
solid answer
~50 s`dbt docs generate` does two things: it compiles the project into `target/manifest.json` (every model, test, source, description and dependency edge) and then queries the warehouse's information schema to build `target/catalog.json` (actual columns, types and table statistics for the relations that exist). `dbt docs serve` spins up a local static site — `index.html` plus those JSON files — on a local port, giving you searchable model pages, the compiled SQL, the tests attached to each column, and an interactive lineage graph you can expand upstream and downstream. Descriptions come from `description:` keys in your schema YAML, optionally through `{{ doc('name') }}` referencing a `{% docs %}` block in a Markdown file so long text is reusable. Most teams run `docs generate` in CI and host the resulting static files rather than expecting analysts to run dbt locally.
code
yaml · 9 linesmodels:
- name: fct_orders
description: One row per completed order, excluding internal test accounts
columns:
- name: order_id
description: Surrogate key, unique per order
tests: [unique, not_null]
- name: status
description: '{{ doc("order_status") }}'go deeper
Know the two commands and that descriptions come from the same YAML where tests are declared — dbt renders documentation you write, it does not invent it.
Explain the two artifacts and where each comes from, and be able to write a docs block and reference it with the doc function.
Talk about hosting: generate in CI after a successful build and publish the static files, plus persist_docs so descriptions reach people who never open the dbt site.
Own whether the dbt catalog is the company's catalog or feeds one, how exposures capture downstream consumers and owners, and how documentation stays reviewed rather than decorative.
## The two artifacts `dbt docs generate` produces two JSON files in `target/`: - **`manifest.json`** — the full compiled representation of the project: every model, snapshot, seed, source, test, exposure and macro, with its config, its compiled SQL, its descriptions, and the dependency edges between them. It is produced by any compiling command, not only docs generate, and it is the file that powers state comparison (`--state`, `state:modified+`) in CI as well as the docs site. - **`catalog.json`** — built by actually querying the warehouse. dbt asks the information schema (or the platform's equivalent metadata views) for the columns, data types and, where available, table statistics of the relations in your project. This is why the command needs a live warehouse connection and why the catalog reflects what physically exists, not what your YAML claims exists. Comparing the two is what powers the site's ability to show you a column that exists in the warehouse but is undocumented, or a documented column that no longer exists. ## The site `dbt docs serve` starts a small local web server on a local port and opens the static site: `index.html` bundled with dbt, reading those two JSON files. What you get per node is the description, the columns with their types and descriptions, the tests attached to each column, the raw and compiled SQL, and the upstream/downstream dependencies. The lineage graph is the part people actually use — a clickable DAG you can filter with dbt's own selection syntax to answer "what breaks if I change this staging model?". Because it is static files, the normal production pattern is to run `dbt docs generate` in CI after a successful build and publish `target/index.html`, `manifest.json` and `catalog.json` to any static host or internal site. Analysts get a link; nobody installs dbt. ## Where the text comes from Descriptions live in the same YAML as tests: ```yaml models: - name: fct_orders description: One row per completed order, excluding test accounts columns: - name: status description: '{{ doc("order_status") }}' ``` For anything longer than a line, or any text repeated across models, use a **docs block**. Create a Markdown file anywhere under the models path: ```markdown {% docs order_status %} The lifecycle state of an order. | value | meaning | |-------|---------| | placed | created but not fulfilled | | shipped | handed to the carrier | | returned | received back | {% enddocs %} ``` and reference it with `{{ doc('order_status') }}`. The Markdown renders on the docs site, and defining the enumeration once means the seven models carrying a `status` column cannot drift apart. ## persist_docs Docs that live only in a dbt site are invisible to someone browsing the warehouse in a SQL client or a BI tool. `persist_docs` fixes that by pushing your descriptions into the warehouse's own table and column comments: ```yaml models: +persist_docs: relation: true columns: true ``` After that, `COMMENT ON` statements accompany the model build, and the description surfaces wherever the platform's metadata surfaces — including many BI tools and catalog products that read it. Support varies by adapter, and it adds statements to each build, but for a team whose analysts live in the warehouse UI it is the highest-leverage documentation config in dbt. ## Exposures and the boundary of the graph One more docs-relevant object: an **exposure**, declared in YAML, represents a downstream consumer — a dashboard, an ML model, an application — with its owner and the models it depends on. Exposures appear in the lineage graph and are selectable (`dbt build --select +exposure:weekly_kpis`), which turns "which dashboards break if this model changes" from tribal knowledge into a query. ## Honest limits dbt docs describes *your dbt project*. It does not discover tables dbt does not manage, it has no notion of who queried what, and the descriptions are only as current as the pull requests that touch them. It is a project catalog with lineage, not an enterprise data catalog, and a candidate who says that plainly — while also saying that a generated, version-controlled, review-gated catalog beats a hand-maintained wiki that is wrong within a month — is giving the answer an interviewer wants.
- Why does dbt docs generate need a warehouse connection?Because catalog.json is built by querying the platform's metadata views for the real columns, types and statistics of the relations in your project. manifest.json comes from parsing and compiling the project files, but the catalog half reflects what physically exists in the warehouse, so it cannot be produced offline.
- How do you reuse one long column description across several models?Write a `{% docs some_name %} ... {% enddocs %}` block in a Markdown file under the models path, then set the column's description to `'{{ doc("some_name") }}'` in each model's YAML. The Markdown renders on the docs site, and defining the text once stops the copies drifting apart.
- What does persist_docs do that the docs site does not?It emits COMMENT statements so your descriptions become the warehouse's own table and column comments. That makes them visible to anyone in a SQL client, BI tool or external catalog that reads platform metadata, rather than only to people who open the dbt docs site. Adapter support varies.
- What is an exposure and why does it matter for documentation?A YAML-declared node representing a downstream consumer — a dashboard, an ML job, an app — with an owner and the models it depends on. It appears in the lineage graph and is selectable, so "which dashboards does this model feed" becomes a query rather than tribal knowledge.
saying these in an interview costs you the question
- Thinks dbt writes the descriptions for you automatically
- Says catalog.json can be produced without a warehouse connection
- Believes docs serve hosts the site for the whole company
- Confuses dbt docs with a full enterprise data catalog
- Assumes descriptions appear in the warehouse without persist_docs