How do you stop a failed data-quality check from publishing bad rows to downstream pipeline consumers?
answer
- check before consumers can read it
- stage, assert, then swap
- the verdict must be a dependency
- stale beats wrong — sometimes
- not every deviation deserves a halt
basics
~20 sWrite to a staging location first, run the checks against it, and publish only if they pass — the write-audit-publish pattern. Make the check a task that downstream work depends on, so a failure leaves the previous good version in place and stops dependent tasks from running.
solid answer
~50 sChecking after publication is too late, so restructure the write. **Write-audit-publish** does three steps: write the new data to a staging table, partition or branch that no consumer reads; **audit** it with the quality assertions; then **publish** atomically — swap the partition, repoint the view, rename the table, or merge the branch. The gate has to be part of the dependency graph, not a dashboard. The audit step is a task; downstream tasks depend on it; a failed assertion fails that task and leaves downstream work unstarted, with the last good published version still serving consumers. That is a **circuit breaker**: you deliberately choose staleness over wrongness. Which is correct depends on the consumer, so tier the checks — blocking severity for the assertions that mean the data is unusable (primary key not unique, volume collapsed, a required column all null), warning severity for the rest, so the pipeline does not halt on a cosmetic deviation at 3 a.m.
code
sql · 14 lines-- 1. WRITE to a staging relation no consumer reads
create or replace table stg.orders_daily__new as
select ... from raw.orders where order_date = date '2026-08-20';
-- 2. AUDIT: each assertion must return zero rows
select order_id from stg.orders_daily__new
group by order_id having count(*) > 1; -- key uniqueness
select 1 from stg.orders_daily__new
having count(*) < 0.6 * 1204388; -- volume floor vs baseline
-- 3. PUBLISH atomically once the audits pass
alter view mart.orders_daily set definition
as select * from stg.orders_daily__new;go deeper
Know that a data-quality check is only a gate if something downstream depends on its result, and that checks belong before consumers can read the output.
Be able to describe write-audit-publish step by step, and name the atomic publish mechanisms — partition swap, view repoint, rename, branch merge.
Show the operational judgement: severity tiering, the freshness-versus-correctness trade per consumer, blast radius of a gate on a shared table, and the documented unblock path.
Own the platform policy — which datasets get blocking gates at all, who is accountable when one halts a hundred pipelines, and how overrides stay auditable rather than becoming routine.
## The failure being prevented A transformation writes directly into the table consumers read. A quality check runs afterwards and fails. By then the dashboards have already refreshed, a reverse-ETL job has already pushed the rows into the CRM, and a downstream model has already consumed them. The alert is real but useless — the damage propagated before anyone could act on it, and cleanup means recomputing every downstream artefact. The fix is structural, not procedural: make publication depend on the check passing. ## Write-audit-publish **WAP** is the canonical pattern: 1. **Write** the new data somewhere no consumer reads — a staging table, a shadow partition, a temporary schema, or a branch in a catalog that supports branching. 2. **Audit** it: run the assertions against the staged data. Uniqueness of the key, expected row-count band, null and range constraints on required fields, referential checks against a dimension, comparison against yesterday's totals. 3. **Publish** with a cheap atomic operation: swap the partition into the target table, `alter view` to point at the new relation, rename tables, or merge the branch. Consumers move from the old version to the new one in one step with no window where they see partial or unaudited data. The key property is that the audit runs on the *real, final* output — not on a sample, not on the source, not on a reimplementation of the logic — but before anyone can read it. ## Wiring the gate into the pipeline Express the audit as its own task and make the publish step, and every downstream task, depend on it. Then the orchestrator's ordinary failure semantics do the work: a failed assertion fails the audit task, the publish never runs, dependent tasks stay unstarted, and consumers continue reading the last known-good version. This is why a quality check that only writes a metric to a dashboard is not a gate — nothing consumes its verdict. A gate should also emit its measurements as run metadata whether it passes or fails. That builds the history the thresholds compare against, and lets you see a metric drifting toward its bound before it trips. ## Blocking versus warning: the real judgement Blocking everything on every deviation produces an unusable platform: at 3 a.m. a marginal null-rate wobble halts a mart that would have been perfectly usable, someone is paged, and within a month the team disables the gate entirely. Blocking nothing produces the original problem. Tier the assertions by severity: - **Blocking (error)** — the data is unusable if this fails: duplicate primary keys, volume outside a wide band, a required column entirely null, a foreign key with unmatched rows, a total that disagrees with the source system. - **Warning** — suspicious but consumable: a distribution shift, a new category value, a modest freshness slip. Publish, record, notify the owner. Tier by dataset too. A regulatory report or a billing feed should stop rather than publish something wrong. An exploratory table should not wake anybody. ## Circuit breaking and the cost of stopping Stopping publication is a **circuit breaker**: it trades freshness for correctness. That trade is not free, and it is not always right. Serving yesterday's marketing dashboard is fine; serving yesterday's inventory levels to a system that decides what to reorder may be worse than serving today's slightly odd numbers. Decide per consumer, and make the decision visible — the dataset should carry a state (`stale, last good version from <time>, blocked by check X`) rather than silently looking normal. Blast radius matters as well. A gate on a widely-consumed staging table halts a hundred downstream pipelines; that is the point, but it means the on-call story, the ownership and the unblock path must exist before you turn it on. ## Recovery A blocked pipeline needs a defined way out, or the gate becomes an outage generator: - **Diagnose from the staged data** — it is still there, which is one of WAP's underrated benefits: the bad output is available for inspection instead of having overwritten the good one. - **Fix and rerun** when the cause is upstream or a code bug. - **Override deliberately** when the deviation is genuine — a real business event, a one-off migration. The override should be an auditable action (an approval, a recorded flag for that run), never a quiet edit to the threshold. - **Notify consumers** using downstream lineage, so the people reading the stale dataset learn it is stale from you rather than from a wrong decision. ## What to say in an interview Name the pattern, place the check between write and publish, make it a dependency rather than a dashboard, explain the severity tiering and the freshness-versus-correctness trade, and finish with the unblock path. Candidates who have actually run one of these always mention the override procedure and the alert-fatigue failure mode, because both bite within the first month.
- Why is running the quality check after writing into the consumer-facing table not equivalent?Because propagation has already started. Dashboards refresh, downstream models consume, reverse-ETL pushes rows to operational systems — all before the alert fires. Cleanup then means recomputing every downstream artefact rather than simply not publishing. Staging the write moves the check in front of every consumer at the cost of one extra atomic operation.
- How do you decide which assertions block publication and which only warn?Block when the data would be unusable or actively harmful: duplicate keys, volume collapse, a required field entirely null, totals disagreeing with the source. Warn on suspicion — distribution drift, a new category, a small freshness slip. Tier by consumer too: regulated or billing outputs stop, exploratory tables do not page anyone.
- What is the cost of a circuit breaker that halts a widely-consumed staging table?Everything below it stops, which is the intent but also the risk: one strict threshold can idle a hundred pipelines. Before enabling it you need an owner, an on-call path, a documented override and a way for downstream consumers to see that their dataset is intentionally stale rather than quietly wrong.
- How should an override be handled when the deviation turns out to be a genuine business event?As an explicit, auditable action for that run — an approval or a recorded acknowledgement — not a quiet loosening of the threshold. Silently widening bounds is how gates decay into decoration. If the event represents a new normal, change the baseline deliberately, in code review, with the reason recorded.
It is the same idea as a build that will not deploy until its tests pass: the artefact exists, but nothing is promoted to production until the assertions are green.
saying these in an interview costs you the question
- Runs the quality check after publishing and calls it a gate
- Reports check results to a dashboard that nothing depends on
- Blocks on every deviation, then gets the gate disabled after alert fatigue
- Assumes stale data is always safer than imperfect data
- Has no documented way to override or unblock a halted pipeline