How do you stop dbt models from building on a stale source table?
answer
- the run does not look at freshness
- two places to put the gate
- before the run, or inside the graph
- a failing source test skips its dependents
basics
~20 sA dbt run never checks freshness, so you must gate it: run dbt source freshness as a separate orchestrator step that fails the pipeline on error, or put a recency test on the source so dbt build skips every downstream model when it fails.
solid answer
~40 s`dbt run` and `dbt build` do not evaluate freshness thresholds — `dbt source freshness` is its own command — so a dead loader produces a perfectly green transform run over yesterday's data. There are two blocking mechanisms. First, run the freshness command as a pre-flight task in the orchestrator; it exits non-zero when a source reaches error state, so the transform task never starts. Second, put a data test on the source that asserts recency, and let `dbt build` do the work: because `dbt build` runs each node's tests before its dependents, a failing source test **skips** every downstream model rather than building it. The first gives you a clean pipeline-level gate and an artifact for dashboards; the second gives you per-source granularity, so a stale marketing feed stops only the marketing marts.
code
bash · 5 lines# pre-flight gate: non-zero exit on error-level staleness
dbt source freshness
# only reached when the gate passes
dbt buildgo deeper
Know that freshness is checked by its own command and that a normal dbt run happily transforms stale data. Recall that source tests exist alongside model tests.
Explain the ordering guarantee that makes this work: dbt build runs a node's tests before its dependents, so an error-severity source test skips the models beneath it rather than building them.
Present both mechanisms with their tradeoffs, choose thresholds that will not be muted, and describe what consumers see when part of the warehouse is skipped and part is refreshed.
Own the policy across teams: which sources carry blocking SLAs versus advisory ones, who is paged, how partial refreshes are signalled to downstream consumers, and how you avoid the alert fatigue that quietly disables the gate.
## Why the run is green The surprise most teams hit is that dbt's transform commands have no opinion about how old the raw data is. `dbt run` compiles your models and executes them against whatever is in the source tables. If the loader died at 2am, the models rebuild flawlessly over stale rows and the dashboard shows yesterday's numbers with today's refresh timestamp — arguably worse than an outright failure, because nothing looks broken. Freshness lives in a separate command, `dbt source freshness`, which reads the `freshness:` blocks in your source YAML and reports pass, warn or error per table. Nothing runs it for you. ## Mechanism one: gate in the orchestrator The standard pattern is a pre-flight task: ```bash dbt source freshness # fails the task when any source is in error state dbt build # only reached if the gate passed ``` The freshness command exits non-zero when a source breaches its `error_after` threshold, so the orchestrator marks the task failed and the downstream transform task never starts. This is coarse — one stale source stops everything — which is exactly right when your sources are mutually dependent and wrong when they are not. A refinement is to run the gate but let it only *warn*, while alerting on the `sources.json` artifact it writes. That keeps the pipeline moving for tolerable lateness and pages a human for real breakage. Which you choose is a policy decision, and interviewers are listening for the fact that you made one deliberately. ## Mechanism two: a test on the source inside dbt build `dbt build` runs the project in DAG order and executes each node's tests **before** the nodes that depend on it. When a test fails with error severity, dbt marks the dependent nodes **skipped** rather than building them. Put a recency assertion on the source and you get selective, dependency-aware blocking for free: ```yaml sources: - name: jaffle_shop tables: - name: orders data_tests: - dbt_utils.recency: datepart: hour field: _etl_loaded_at interval: 24 ``` Now only the models downstream of `jaffle_shop.orders` are skipped; the finance marts fed by a different, healthy source still build. A singular test — a SQL file that returns rows when the condition is violated — does the same job without a package dependency. Severity is the knob here. A test configured with `severity: warn` reports and continues; error severity is what actually stops dependents. Some teams set `warn_if` and `error_if` thresholds so a small breach is visible without halting the warehouse. ## Choosing between them The orchestrator gate is a *pipeline* statement: nothing runs on stale data, period. It is easy to reason about, produces one clear alert, and is the right default for a small project with one loader. The in-build test is a *graph* statement: each subgraph fails independently. It scales better as the project grows and multiple ingestion pipelines with different SLAs feed the same warehouse. Its cost is that skipped models leave a partially refreshed warehouse — some marts current, some not — so consumers need to be able to tell which is which, usually via a freshness or last-updated column on published tables. Many mature setups run both: the freshness command for observability and dashboards, and source tests for selective blocking. ## The related selector The inverse problem — not blocking on stale data, but avoiding pointless rebuilds of models whose sources have not changed — is solved by state comparison against a previous `sources.json`. Selecting `source_status:fresher+` builds only the models downstream of sources that received new data since that artifact was produced. It is a cost optimisation, not a safety gate, and confusing the two in an interview is a tell. ## What weak answers look like "We check the dashboard" is not a gate. "We put a `not_null` test on the source" does not detect staleness — old rows are perfectly non-null. "We rely on the loader to alert us" pushes the guarantee onto a system that is, by hypothesis, the one that failed. And "we fail everything on any warning" is how alert fatigue starts: within a month someone adds `--warn-error` exceptions or people stop reading the channel. ## The judgment to show Name both mechanisms, state which threshold blocks and which only warns, and be explicit about the blast-radius tradeoff: pipeline-wide halt versus subgraph skip, and what consumers see in each case.
- What is the practical difference in blast radius between the two approaches?The orchestrator gate halts the entire transform, so the warehouse stays wholly on yesterday's state — consistent but entirely stale. The source test skips only the subgraph below the affected source, so some marts refresh and others do not. The second scales better with many independent feeds, but consumers then need a published last-updated signal to tell which tables are current.
- Does a not_null or unique test on a source catch staleness?No. Those assertions hold perfectly well on data that stopped arriving a week ago — old rows are still non-null and still unique. Staleness needs an explicit recency assertion comparing the newest load timestamp against now, either through a recency test from a package or a singular test that returns rows when the lag exceeds your threshold.
- What does selecting source_status:fresher+ do, and is it a safety gate?It compares against a previous sources.json artifact and builds only the models downstream of sources that received new data since then. That is a cost and runtime optimisation — it skips work that would be a no-op — not a guarantee that anything is fresh enough to publish. Blocking on staleness still needs the freshness gate or a source test.
saying these in an interview costs you the question
- Assumes dbt run or dbt build fails automatically on stale sources
- Puts not_null tests on a source and calls it a freshness gate
- Says a failing source test still builds downstream models
- Blocks the whole pipeline on warn-level staleness
- Relies on the loader's own alerting as the only guarantee