skip to content

Each bulk import pipeline on your platform decides for itself whether a rejected record ends the run — what standard would you set?

level: principalimportance: nice to knowfreq 30%

answer

  1. terminality is a claim about the remainder
  2. would the next record succeed?
  3. environment faults stop, record faults do not
  4. rejections need an owner and a threshold
  5. standardise structure, let teams set numbers

basics

~20 s

Define terminality by whether the remainder of the run still has value. Environment-level faults end the run; per-record rejections travel as data with a required destination, a count on every run report, and a rate threshold that stops a run when the input contract has clearly changed.

solid answer

~50 s

The decision is not "is this an error" but "does the rest of this run still have value". Ending a sequence asserts that it does not, so the standard should make that assertion explicit. I would publish a two-channel rule: conditions that make every remaining record unprocessable — the destination refusing writes, credentials rejected, the source truncated — end the run; properties of a single record do not, and ride the sequence as tagged outcomes. Then I would attach the obligations that keep the second channel honest: every rejection has a named destination and an owner, every run reports accepted, rejected and aborted separately, a rejection-rate threshold aborts and alerts, and reruns are safe over records already written. Both extremes fail — always-terminal gives brittle jobs that never finish, never-terminal gives clean-looking runs that quietly drop data.

go deeper

for a junior

The takeaway is that ending a run is a decision, not an accident: it says the rest of the work is not worth doing, and that is only true for some kinds of failure.

for a middle

Be able to sort conditions with one test — would the next record succeed? Environment faults answer no for every record and justify stopping; single-record faults answer yes and belong in the value channel.

for a senior

Show the operational half: rejects need an owner, a queryable destination, counts on the run report and a rate threshold, or the policy trades loud aborts for quiet data loss.

for a principal

Own the trade-off across teams. Standardise the run-outcome vocabulary and the obligations, let each feed set its own numbers from its own history, put the shape in shared scaffolding, and bring abort and unread-rejects data to justify the change.

## The question behind the standard Ending a sequence is a claim about the **remainder**: it says the rest of this run is not worth producing. Most teams never state that claim, so they end up with two accidental conventions — some pipelines end on anything unexpected and never finish a large file, others end on nothing and report success while dropping records. A standard is worth setting because the underlying decision is not a matter of taste; it follows from a property of the condition. ## A split that holds up 1. **Environment-level conditions end the run.** The destination refuses writes, credentials are rejected, the source is unreadable or truncated. Every remaining record would meet the same wall, so the remainder has no value and stopping is honest. 2. **Per-record conditions do not.** A malformed field, a value outside the allowed range, a reference to something that does not exist. The next record will probably be fine, so the run keeps value and the outcome belongs in the value channel as a tagged rejection. 3. **Aggregate conditions end the run deliberately.** A rejection rate far above the historical norm usually means the wrong file, a changed upstream contract, or a broken producer. This is where terminality genuinely belongs for rejections — at a threshold you chose, not at the first bad record you happened to meet. The test that decides which bucket a condition falls in is a single question: **would the next record succeed?** Environment faults answer no for every record; per-record faults answer yes. ## What a carried rejection must come with Moving rejections out of the ending is only safe if they land somewhere. Require, as part of the standard: - **A named destination** for rejected records, with the reason attached — not a log line, something an operator can query and re-submit from. - **A named owner** who is expected to look at it. A rejects destination with no owner is a slow leak. - **A distinct run outcome vocabulary**: *completed*, *completed with rejections*, *aborted*. Collapsing the first two into "success" is how data loss becomes invisible. - **Counts on every run report**: accepted, rejected by reason, and the rejection rate against the historical baseline. - **A threshold that aborts and alerts**, tuned per pipeline from its own history. - **Rerun safety**, since any run that aborts mid-flight leaves partial work durable. ## The two ways the standard fails | | everything terminal | nothing terminal | |---|---|---| | symptom | large runs rarely finish; each attempt stops at a different record | runs always report success | | what is lost | the night's work, repeatedly | records, quietly, for as long as nobody reads the rejects | | how it is discovered | immediately and loudly | weeks later, by someone downstream | | cost of discovery | operator time | trust in the data and a backfill | | the error in the model | treats every anomaly as a claim about the remainder | never makes the claim even when it is true | The asymmetry is the argument for making the second channel carry obligations: an over-terminal pipeline is annoying, an under-terminal one is dangerous, and the only thing standing between them is whether rejections have an owner and a threshold. ## What to standardise and what to leave alone Standardise the **vocabulary and the obligations**: the three run outcomes, the required rejects destination, the required counts, the existence of a threshold, and rerun safety. Leave the **numbers** to teams — the right threshold for a feed that is normally flawless is not the right one for a feed from a partner that has always been noisy. A standard that fixes the numbers gets ignored or gamed; one that fixes the structure and demands each team set and justify its own numbers survives. Make it real in shared pipeline scaffolding rather than in a document. If the default shape a team starts from already has both channels and refuses to report plain success when records were rejected, the standard holds without enforcement. A written convention alone drifts within two quarters. ## Deciding with evidence Before proposing any of this, measure two things: how often runs abort and at what position, and how many rejected records sat unread. The first tells you how much work the current defaults are throwing away; the second tells you how much data the alternative is already losing. Bring those numbers, and the standard argues for itself.

  • What makes a rejections-as-data policy fail in practice?
    A rejects destination that nobody owns. The run reports success, the records sit unread, and the loss surfaces weeks later as a gap someone downstream notices. The policy only works when the rejects have an owner, a visible count and a rate threshold that can stop a run.
  • How would you pick the rejection-rate threshold for a feed that is normally almost clean?
    From its own history rather than a platform-wide number. If a feed normally rejects well under one per cent, a run rejecting ten per cent is evidence the file or the upstream contract changed, and continuing pollutes the destination. Set it where the distribution actually breaks, and review it.
  • Should this live in a written convention or in code?
    In shared pipeline scaffolding that makes both channels explicit and refuses to report plain success when records were rejected, backed by a short written rule. Structure that is inherited holds; a document alone drifts as teams and deadlines change.

saying these in an interview costs you the question

  • Declares one blanket policy without weighing the cost of each direction
  • Routes rejections to a destination nobody owns or reads
  • Reports a run as successful while records were silently dropped
  • Treats an unreachable destination as a per-record rejection
  • Sets platform-wide thresholds ignoring each feed's own history
  • Proposes the standard without measuring current aborts and unread rejects