dbt
Transformation as version-controlled SQL: models referencing each other, materializations, tests, snapshots and Jinja macros, all executed inside the warehouse. Interviewers ask about dbt because it is how analytics engineering brought software practice to SQL.
on this pageshowhide
explore
- Models and Refs6 questions
- Materializations6 questions
- Sources and Seeds7 questions
- Tests and Documentation7 questions
- Jinja and Macros6 questions
- Snapshots and SCD6 questions
questions
page 2 of 2How do you decide how much Jinja templating a dbt project should allow in its models?
basics
~20 sAllow Jinja that removes duplication without hiding what a model does: shared expressions, environment switches, config. Ban templating that decides a model's grain, joins or output columns, because reviewers then read a template while the warehouse runs something else.
When should reference data live as a dbt seed rather than a pipeline-loaded table?
basics
~20 sSeed it when the analytics team is the system of record and the data is small, slow-moving and non-sensitive, so a pull request is the right change process. Ingest it when an upstream system already owns the data or it changes often.
In dbt, what does the ephemeral materialization do to a model's compiled SQL?
basics
~20 sAn ephemeral dbt model creates no object in the warehouse. Its SQL is interpolated as a common table expression into every downstream model that refs it, so the logic is reusable in dbt but not queryable outside it.
In a dbt snapshot table, what does the dbt_scd_id column identify?
basics
~20 sdbt_scd_id is a generated hash that uniquely identifies one version row in a dbt snapshot — the entity's unique_key combined with the timestamp that version starts at. It is dbt's internal row identity, not a business key.
A dbt seed of ZIP codes loads as integers and loses leading zeros — how do you fix it?
basics
~20 sdbt infers seed column types from the CSV contents, so an all-digit column becomes numeric. Pin it with a column_types config naming the column as a string type, then run dbt seed with full refresh so the table is dropped and recreated.
In dbt, how do you write a custom generic test that can be reused from schema YAML?
basics
~20 sDefine a Jinja block with the test tag taking model and column_name arguments, save it under tests/generic or macros, and return the failing rows. Any model's YAML can then declare it by name and pass extra arguments.
In dbt, how does adapter.dispatch let one macro work across Snowflake and BigQuery?
basics
~10 sadapter.dispatch resolves a macro name to an adapter-prefixed implementation at compile time: on Snowflake it looks for snowflake__my_macro and falls back to default__my_macro. You write one entry point plus one implementation per warehouse.
In dbt, what do model access levels and groups control about who can ref() a model?
basics
~20 sAccess marks a model private, protected or public. Private models can only be referenced from within their own group, protected ones from within the project, and public ones are the declared interface other projects may reference — turning the graph into an ownership boundary.
showing 31–38 of 38