Which ingestion validation failures should fail the whole load instead of quarantining rows?
answer
- ask what the failure implies about the rest
- one row, or the whole delivery?
- counts and totals are batch-scope
- a huge reject rate means the source changed
- flagging keeps the measure in the total
basics
~20 sFail the load when the defect makes the whole batch untrustworthy: an unrecognisable schema, a row count or control total that disagrees with the manifest, duplicate keys that would corrupt a merge, or a reject rate far above baseline. Independent per-row defects belong in quarantine.
solid answer
~50 sSplit checks by **scope**, not by severity. *Batch-scope* checks describe the delivery as a whole — does the header match the contract, does the row count match the manifest, do control totals reconcile, is the reject rate inside its normal band, are the primary keys unique. A failure there means you cannot trust any row, so the load must abort before it writes anything; loading half a corrupt file is worse than loading none of it. *Row-scope* checks describe one record independently — a malformed email, an unparseable optional date, an unknown enum value. Those failures are isolated, so the row goes to quarantine and the batch proceeds. A third mode, **pass-through-and-flag**, loads the suspect row into the target with a validity flag: use it when excluding the row distorts an aggregate more than including a dubious one does. The trap to name is the silent one — quarantining rows that carry measures shrinks the day's totals unless completeness is reported alongside.
code
yaml · 19 lineschecks:
- name: header_matches_contract
scope: batch
on_failure: fail_load
- name: row_count_matches_manifest
scope: batch
on_failure: fail_load
- name: order_id_unique_in_batch
scope: batch
on_failure: fail_load # a merge on a duplicated key is non-deterministic
- name: reject_rate_within_budget
scope: batch
on_failure: fail_load
- name: email_format_valid
scope: row
on_failure: quarantine
- name: currency_code_known
scope: row
on_failure: flag_and_load # keeps the amount in the day's totalgo deeper
Learn the split: a defect in one row is that row's problem, a defect in the delivery is everyone's. Naming row-count and header checks as the ones that stop a load is a solid answer at this level.
Explain the scope axis and give concrete batch-scope examples — manifest counts, control totals, duplicate merge keys, reject rate. Be ready to describe pass-through-and-flag as a third mode and when it applies.
Show the operational consequences: staging-then-promote so an abort leaves nothing half-written, tighter reject thresholds on measure-bearing tables, and completeness metrics published with every run.
Own it as declared policy rather than code branches — each check carries its scope and on-failure action, source owners can argue about a specific rule, and the same defect always yields the same outcome across pipelines.
## Scope is the deciding axis, not severity Engineers reach for "how bad is this?" and get lost, because a null in a nullable column and a null in a join key feel equally bad in isolation. The useful question is **what does this failure tell me about the rest of the batch?** If the answer is "nothing — the other rows are unaffected", the failure is row-scope and quarantine is correct. If the answer is "this means the delivery itself is wrong", it is batch-scope and the load must not proceed. ## Batch-scope failures: abort before writing These are the checks that run once per delivery, ideally before any target write: - **Structure.** The header, column set or schema does not match what the contract declares. If a source silently reorders columns, per-row validation may pass while every value lands in the wrong field. This is the failure that looks fine and is catastrophic. - **Completeness against a manifest.** The delivery declares 4.2 million rows and the file holds 3.1 million, or a partition's expected files are missing. Loading it produces a batch that is quietly short. - **Control totals.** The source declares a sum of amounts; your parsed sum disagrees. Something was lost or misread and you cannot say what. - **Key integrity for the write pattern.** If the load performs a merge keyed on a business key, duplicate keys in the incoming set make the merge non-deterministic — many engines error, and the ones that do not pick arbitrarily. This is a correctness failure of the write, not of one row. - **Reject rate outside its band.** A batch where 40% of rows fail a rule almost never means the data got worse; it means the source changed. Halting turns an invisible degradation into a visible incident. - **Freshness or identity of the delivery.** Yesterday's file re-dropped, or a file whose partition prefix says a different date than its contents. The operational rule that goes with these: validate the batch *before* the target write, or write to a staging area first and promote atomically, so a batch-scope abort leaves the target untouched rather than half-loaded. ## Row-scope failures: quarantine and keep going These are defects contained inside one record and independent across records — an invalid email format, a free-text field over its declared length, a timestamp that will not parse, a categorical value not in the allowed set, a reference key that resolves to nothing. One customer's mistyped postcode says nothing about the next customer's. Halting the batch for these is how pipelines become brittle and how tired engineers learn to delete rows from staging at 03:00. ## The third mode: pass through and flag Sometimes removing a row is more damaging than admitting it. A revenue fact whose currency code is unrecognised still carries a real amount; quarantining it makes the day's revenue silently low, and an unexplained dip is harder to notice than an explicit flag. The pattern is to load the row with a boolean or a quality score — `is_valid`, `dq_status`, a rejected-rules array — so completeness-sensitive consumers see it and cleanliness-sensitive consumers filter it. Dimension loads use the same idea with an "unknown member" row so facts never lose their join. Be deliberate: flagging is only safe when downstream consumers actually respect the flag. If the marts select `*` and nobody filters, flagging is a way of loading bad data with a note attached. ## The aggregate trap The defect this decision hides is **silently shifting a measure**. Quarantining 2% of order rows means every downstream sum, average and conversion rate is off by roughly 2% and every dashboard still renders confidently. The mitigation is not to stop quarantining — it is to publish completeness alongside the data: rows read, rows loaded, rows rejected, per run, and to fail the load when the rejected share of a *measure-bearing* table crosses its threshold. Rows that carry no measures (a lookup, an event attribute) tolerate a far looser threshold than rows that carry money. ## Making the policy explicit The strongest answer in an interview is that the mode is a property of the check, declared where the check is declared, rather than a decision made in the exception handler. Each check carries a scope (batch or row) and an on-failure action (fail the load, quarantine, flag and load). That makes the policy reviewable, lets a source owner argue about a specific rule rather than about the pipeline's temperament, and means the same defect always produces the same outcome regardless of which engineer wrote the branch.
- Why can quarantining a small share of rows still corrupt a dashboard?Because the rows carried measures. Quarantining 2% of orders makes every downstream sum and rate about 2% low, and the dashboard renders that confidently with no visible defect. The fix is to publish rows read, loaded and rejected per run, and to set a much tighter reject threshold on measure-bearing tables than on lookups.
- How do you make a batch-scope failure leave the target untouched?Validate before the target write, or write into a staging table or a new partition first and promote atomically only after the batch checks pass. Streaming into the target row by row means an abort halfway leaves a partially applied batch, which is exactly the state batch-scope checks exist to prevent.
- When is pass-through-and-flag the wrong choice?When downstream consumers do not honour the flag. If the marts select everything and nobody filters on the validity column, flagging simply loads bad data with a note attached. Flagging only works where the filter is enforced in a shared model layer or the flag is part of the published contract.
saying these in an interview costs you the question
- Deciding fail-vs-quarantine by gut severity instead of scope
- Quarantining measure-bearing rows without reporting completeness
- Letting a batch abort halfway through writing the target
- Treating a 40% reject rate as bad data rather than a changed source
- Flagging rows as suspect when no consumer ever filters on the flag