In dbt, what do model access levels and groups control about who can ref() a model?
answer
- not every model is meant to be reused
- a boundary the parser enforces
- who owns it, and who may read it
- the public ones need a promised shape
basics
~20 sAccess marks a model private, protected or public. Private models can only be referenced from within their own group, protected ones from within the project, and public ones are the declared interface other projects may reference — turning the graph into an ownership boundary.
solid answer
~40 sBy default a dbt model is `protected`: anything inside the project can `ref()` it. Setting `access: private` restricts references to models in the same **group** — a named collection of models with an owner declared in YAML — so a team's internal intermediate models stop being quietly reused by other squads. Setting `access: public` declares the model a stable interface; in a multi-project setup, only public models can be referenced from another project, using the two-argument `ref('other_project', 'model_name')` form. The point is governance: without it, every model in a large repo is equally referenceable, so today's throwaway intermediate model becomes tomorrow's undeclared dependency. Public models are usually paired with a **contract** that enforces column names and types, and with **versions**, so the interface can change without silently breaking consumers.
code
yaml · 14 linesgroups:
- name: finance
owner:
name: Finance Analytics
email: [email protected]
models:
- name: int_finance__ledger_joined
group: finance
access: private
- name: fct_revenue
group: finance
access: publicgo deeper
Know that this exists and that the default lets anything in the project reference anything else. You will not be expected to design the boundaries yet.
Be able to name the three access levels, say that private is scoped to a group, and explain that enforcement happens at parse time rather than at run time.
Explain why an undeclared internal model becoming load-bearing is the real failure, and pair public access with contracts and versions so the published shape can change safely.
Own the call: one project with groups and access versus multiple projects with cross-project references, judged on release cadence, permission isolation and coordination cost — and be explicit that the multi-project workflow depends on the commercial platform.
## The problem this solves In a small dbt project every model is referenceable by every other model, and that is fine. At a few hundred models with several teams contributing, it stops being fine. Someone's intermediate model — written as a private stepping stone, never tested at an interface level, expected to change shape freely — gets referenced by another team because it happened to have the columns they needed. Now it is load-bearing and nobody agreed to that. The project's DAG has become an undocumented API with no owner and no compatibility promise. Model **access** and **groups** exist to make that boundary explicit and enforced at parse time rather than by convention. ## Groups A **group** is a named collection of resources with a declared owner, defined in YAML: ```yaml groups: - name: finance owner: name: Finance Analytics email: [email protected] ``` Models are assigned to a group by config, individually or by directory in `dbt_project.yml`. On its own a group is documentation plus an ownership record; combined with access it becomes a boundary. ## The three access levels - **`protected`** — the default. The model can be referenced by anything in the same project. This is the historical behaviour, so an existing project keeps working unchanged when you upgrade. - **`private`** — the model can only be referenced by models in the **same group**. A `ref()` from outside fails at parse time with a clear error, before any SQL runs. This is the level for staging and intermediate models a team wants to keep refactorable. - **`public`** — the model is declared a stable interface. It is referenceable widely, and in a multi-project setup it is the only access level another project's `ref()` can reach. The enforcement is static: dbt refuses to build a project containing an illegal reference, so the boundary cannot erode silently the way a review-time convention does. ## Contracts and versions on the public surface Declaring a model public without pinning its shape only moves the problem. Two companion features close it: - A **contract** on a model asserts the column names and data types it will produce, enforced when the model builds. If the SQL stops producing a contracted column, the build fails at the producer rather than in a consumer's dashboard days later. - **Versions** let a public model exist as `v1` and `v2` simultaneously, with a declared latest. Consumers reference a specific version, migrate when they are ready, and the old version is deprecated on a schedule rather than deleted from under them. Together these give the thing that was missing: a published model with a shape guarantee and a migration path. ## One project or many The strategic question a lead actually owns is whether teams share one repository or split into several projects that reference each other's public models. **One project with groups and access** keeps everything in a single graph: one command builds the world, cross-team lineage is trivially complete, refactors that span teams are one pull request. The costs are a monorepo's usual ones — CI runtime, merge contention, a `dbt_project.yml` that many people edit — mitigated by state-based selection so a pull request only rebuilds what it touched. **Multiple projects** give each team its own repository, release cadence, CI and permissions, with cross-project `ref('project', 'model')` reaching only public models. The cost is a coordination surface: a consumer builds against the producer's *last successful production run*, so a breaking change is discovered across a repo boundary, and lineage now spans systems that must agree on state artifacts. In dbt this multi-project workflow is a dbt Cloud capability rather than something the open-source runner does on its own — worth being precise about in an interview. The honest recommendation is usually: start as one project, introduce groups and access as soon as more than one team contributes, and split into separate projects only when release-cadence independence or permission isolation genuinely justifies the coordination cost — not merely because the repository feels large. ## What interviewers are probing This is a judgment question, and there is no single right answer. What distinguishes a strong response is that it treats the model graph as an **API surface with owners** rather than a pile of SQL: which models are promised to outsiders, what shape they promise, how that promise changes over time, and who is paged when it breaks. A candidate who only answers "private, protected, public" has the trivia; one who explains why an undeclared internal model becoming load-bearing is the actual failure has the judgment.
- What does a model contract add on top of declaring a model public in dbt?Access says who may reference the model; a contract says what shape it guarantees. The contract asserts column names and types and is enforced when the model builds, so a change that drops or retypes a contracted column fails at the producer. Without it, a public model's promise is only a naming convention.
- When would you split a dbt monorepo into multiple projects rather than using groups and access?When teams need independent release cadences, separate CI ownership, or genuinely separated warehouse permissions — not merely because the repo is large. Splitting adds a coordination surface: consumers build against the producer's last production run, so breaking changes surface across repository boundaries. Groups and access give most of the isolation with none of that cost.
- What happens if a model outside the group refs a private model in dbt?Parsing fails with an access error and nothing is built. Enforcement is static rather than advisory, which is what makes the boundary hold: an illegal reference cannot be merged and quietly relied upon, so a team can keep refactoring its internal models without discovering hidden consumers later.
saying these in an interview costs you the question
- Thinks access levels control warehouse grants or user permissions
- Assumes private is the default rather than protected
- Believes any model can be referenced across projects
- Treats splitting into multiple projects as strictly better than one
- Declares models public without any contract or versioning plan