skip to content

At what quarantine reject rate should an ingestion pipeline stop rather than keep loading?

level: principalimportance: nice to knowfreq 34%

answer

  1. no universal number exists
  2. baseline the feed before judging it
  3. rate, with a floor for small batches
  4. near-total rejection means a schema change
  5. who is harmed by partial data?

basics

~20 s

There is no universal number. Set the threshold per source and per reason code against that feed's own baseline, as a rate rather than a count, and stop when a partial load would mislead consumers more than a late load would delay them.

solid answer

~60 s

Treat the reject rate as a **circuit breaker with a business-set trip point**, not a constant. Baseline each source and each reason code — a feed that normally rejects a fraction of a percent for bad postcodes and one that normally rejects several percent of optional fields deserve different bands. Alert on **deviation from baseline** rather than an absolute figure, use a **rate** so volume changes do not create false alarms, and add an absolute floor for low-volume feeds where three rejects out of forty is not a signal. Distinguish the shapes: a jump toward total rejection is a schema or contract change, not deteriorating data, and should stop the load; a slow drift upward is a producer problem that warrants a ticket, not a page. Then decide by blast radius — feeds carrying money or feeding regulatory output trip early; exploratory feeds trip late or never. Whatever the number, the breaker needs a documented owner, a manual override, and an alert that reaches the **producer**, because the platform team cannot fix the data.

code

sql · 10 lines
sql
-- reject share for the current run, by source and reason, for the breaker check
SELECT q.source_system,
       q.reject_reason,
       COUNT(*)                                    AS rejected_rows,
       COUNT(*)::numeric / NULLIF(r.rows_read, 0)  AS reject_rate
FROM   quarantine_records q
JOIN   ingest_runs r ON r.run_id = q.run_id
WHERE  q.run_id = :run_id
GROUP  BY q.source_system, q.reject_reason, r.rows_read
ORDER  BY reject_rate DESC;

go deeper

for a junior

Take away that the acceptable reject rate depends on the source and on who uses the data, and that a sudden jump usually means something structural changed upstream rather than the data getting worse.

for a middle

Be able to compute and monitor reject rate by source and reason, explain why a rate beats a count, and say why small batches need an absolute floor before any rate is meaningful.

for a senior

Show operational judgment: baselining a feed, reading spike-versus-drift, tripping the load before a partition is promoted, and keeping the quarantine store's own backlog under alert.

for a principal

Own the tradeoff explicitly — incomplete-and-on-time versus complete-and-late — decide it with consumers rather than in a config constant, name the accountable producer, define the override and audit path, and publish completeness so consumers can judge a partial load themselves.

## Why there is no single number Asking for "the" reject-rate threshold is asking for a constant that spans a payments feed, a clickstream and a scraped third-party CSV. The rate that means *disaster* on the first is *Tuesday* on the third. Anyone who answers with a figure has skipped the actual question, which is: **at what point does continuing to load produce output more harmful than producing nothing?** That is a judgment about consumers, not about data. ## Baseline per source, per reason Start by measuring. Every feed has a normal band of rejects, and it is stable enough to characterise: record reject rate by source and by reason code per run for a few weeks, and take the distribution, not the mean. Then alert on deviation — the run is outside the band this feed has occupied — rather than on a hand-picked constant. Per-reason matters as much as per-source: an overall rate flat at 1% can hide a brand-new reason code that has just taken over the whole of that 1% while the previous cause vanished. **A reason code appearing for the first time deserves attention regardless of rate**, because it means the source is doing something it has never done before. ## Rate, not count, with a floor A count-based threshold fires on Black Friday and never fires on a quiet Sunday, tracking traffic rather than quality. Use a share of rows read. But rates are unstable at small volumes — one reject in twelve rows is 8% and means nothing — so pair the rate with an absolute floor, and suppress alerts below a minimum row count. Conversely, on very large feeds a rate can hide a serious absolute number, so alerting on both shapes and requiring either to trip is the practical compromise. ## Read the shape, not just the level Three signatures, three responses: - **Near-total rejection.** Almost every row failing means structure, not content: a renamed column, a changed format, an expired credential producing an error page in place of data. Stop the load. Continuing produces an empty or wildly incomplete partition that downstream will consume as truth. - **A step change to a new plateau.** A source began emitting a value your rules do not accept — a new enum member, a new region, a wider field. Frequently the rules are wrong, not the data. This is a stop-and-discuss, and it is the argument for versioned contracts with a change process rather than for a tighter breaker. - **A slow drift.** Rejects creep up over weeks. Nothing is on fire; something is rotting. This is a ticket to the producer with evidence, and the metric to watch is the trend, not any single run. ## Blast radius sets the trip point The threshold should be a function of what the data feeds: - **Money, regulatory or externally-published output.** Incomplete is worse than late, almost always. Trip early, block promotion of the partition, and require a human to release it. - **Operational dashboards and internal analytics.** Late data has a real cost too; trip at a clear deviation and prefer loading with a visible completeness figure over halting. - **Exploratory or enrichment feeds.** Loading what parsed is usually fine; log and report, do not stop the world. This is also the honest answer to "why not just always stop?" — halting has a cost that lands on people, and a breaker that trips on ordinary variation gets disabled within a month, after which nothing protects anything. ## The breaker needs governance, not just a number A threshold without an operating model is a page nobody can action. What has to exist alongside it: - **A named owner per feed** — usually the producing team — and alerts routed to them, because the platform team can only stop the pipeline, not fix the values. - **A documented override.** Someone must be able to say "we know, load it anyway, here is why", with that decision recorded against the run so consumers can see it. - **A consumer-visible completeness signal.** Publish rows read, loaded and rejected with the dataset, so a partially-loaded partition is never silently indistinguishable from a whole one. A breaker that trips is visible; a breaker that does not trip must still make the reject share legible. - **Quarantine-store health as its own alert.** Age of the oldest untriaged record and total unresolved volume. A feed can sit under its reject threshold forever while accumulating a growing pile nobody reads. - **Retention and access policy** for the rejected payloads, which are unmasked production data. ## The tradeoff to state plainly The reject threshold encodes a single choice: **incomplete-and-on-time versus complete-and-late**. Different consumers of the same table answer that differently, which is why the strongest version of this design puts the decision in the contract with the consumer rather than in a config constant chosen by whoever built the pipeline — and why the number is reviewed when consumers change, not set once and forgotten.

  • Why is an overall reject rate insufficient even when it looks stable?
    Because the mix can change underneath it. A feed flat at 1% may have swapped one cause for a brand-new reason code that has never appeared before — a structural change hiding inside a stable number. Alert on per-reason rates and on the first appearance of any new reason code, not only on the aggregate.
  • Who should the reject-rate alert page, and why does it matter?
    The producing team that owns the data, not the platform team that owns the pipeline. Platform engineers can only stop or resume the load; they cannot correct values or explain a source change. Routing to the platform team turns every quality incident into a relay and makes producers structurally unaccountable for what they emit.
  • Should a tripped breaker ever be overridable?
    Yes, explicitly and with an audit trail. A breaker with no override is disabled the first time it blocks a genuine business deadline, after which nothing protects anything. Give a named owner the ability to release a run, record the decision and the reason against that run, and surface it to consumers alongside the completeness figures.

saying these in an interview costs you the question

  • Quoting one universal reject percentage for every source
  • Alerting on reject counts so the threshold just tracks traffic volume
  • Treating near-total rejection as bad data rather than a schema change
  • Paging the platform team for defects only the producer can fix
  • Setting a breaker with no override, so it gets switched off entirely

context