skip to content

A four-stage screening chain extracts fields then judges inclusion; how do you stop stage-2 errors poisoning stage 3?

level: seniorimportance: must knowfreq 55%

answer

  1. per-stage accuracy multiplies
  2. silent, plausible, well-formed wrongness
  3. make values quote the source verbatim
  4. let the stage say it does not know
  5. label the golden set per stage

basics

~20 s

Put a check at the seam. Validate the extracted fields against a schema, require each value to appear verbatim in the source, allow the stage to abstain rather than guess, and pass the source forward so the judging stage can disagree with the extraction.

solid answer

~50 s

The danger in a multi-stage chain is not loud failure but plausible failure: an extraction stage returns well-formed fields that are subtly wrong, and every later stage consumes them faithfully, so the final decision is confidently incorrect with no error in the logs. Independent per-stage accuracy also multiplies, so four stages at ninety-five percent each land near eighty-one percent end to end. Contain it at the boundary. Assert the output shape, then add a grounding check that each extracted value appears verbatim in the source text with an offset or quote. Give the stage an explicit abstain option so uncertainty surfaces as a routed exception rather than an invented value. Carry the source alongside the extraction so the judging stage can contradict it. Then hold per-stage golden sets, so when the end-to-end number drops you can say which stage owns it.

go deeper

for a junior

Know that in a multi-step pipeline a mistake early on gets carried forward, and that checking the output of each step before passing it on is what stops it.

for a middle

Explain why per-stage accuracy multiplies and why the dangerous failures are well-formed and plausible. Name concrete checks: schema validation, verbatim grounding of extracted values, an abstain option.

for a senior

Show how you would locate the failing stage in production: per-stage golden labels, correlation ids in logs, isolating a stage by feeding it known-good upstream input, and capped retry with a fallback route.

for a principal

Own the question of where checking effort is spent. Argue that a boundary you cannot write an assertion for should be merged, and set the policy for what happens to records the pipeline is not confident about.

## Why chained errors are worse than single-call errors A literature-screening pipeline that fetches an abstract, extracts structured fields such as population, intervention, comparator, and outcome, judges inclusion against protocol criteria, and drafts a rationale has four places to be wrong. Two properties make this harder than a single wrong answer. The first is arithmetic. If the stages fail independently, end-to-end accuracy is the product of the per-stage accuracies. Four stages at ninety-five percent give roughly eighty-one percent; at ninety percent each, roughly sixty-six. Adding a stage always adds a multiplier below one, which is the quiet cost of every extra seam. The second is worse: the errors are silent. Stage 2 emits a syntactically perfect object whose intervention field says the wrong drug. Stage 3 has no way to know, so it applies the criteria correctly to false facts and returns a well-reasoned exclusion. Stage 4 writes a fluent rationale for a decision that should have gone the other way. Nothing threw, nothing looked odd, and the output is more persuasive than a single confused answer would have been. A reviewer sampling outputs sees clean prose and stops looking. ## The defences, roughly in order of value per unit effort **Schema validation at the seam.** Parse the stage output against a strict schema before the next stage sees it: required fields present, enums within range, types correct. This catches only structural failures, but they are free to catch and would otherwise become downstream nonsense. **Grounding checks.** For anything the stage claims to have taken from a source, verify it is really there. Require the stage to return a verbatim quote or character offsets alongside each field, then check with ordinary string matching that the quote occurs in the source. A fabricated or drifted value fails a deterministic test rather than a judgement call. This is the single highest-value check for extraction stages because it converts a semantic error into a syntactic one. **An explicit abstain path.** Models under-report uncertainty when the output shape has no room for it. Give the stage a legitimate way to say the abstract does not state the comparator, and treat that as a first-class outcome routed to review or to a fallback, not a failure. Without it, uncertainty is expressed as a plausible invention, which is exactly the failure you cannot detect downstream. **Carry the source forward.** The judging stage should see the extracted fields and the abstract, not the fields alone. It costs tokens and it lets a later stage catch what an earlier stage got wrong, which is the only cheap form of redundancy a chain has. Where the source is too large to carry, carry the quoted evidence spans. **Cross-checks between stages.** If stage 3 excludes a record on a criterion, require it to cite which extracted field drove the decision. A decision citing a field that failed its grounding check, or citing no field at all, is a rejectable output. Consistency requirements like this are cheap and catch reasoning that drifted away from the data. **Cheap deterministic checks before expensive model ones.** Date ranges, unit sanity, sample sizes that exceed a plausible bound, an outcome that is not in the protocol's controlled vocabulary. A dozen assertions written in an afternoon typically catch more real defects than another round of prompt tuning. ## Making failures attributable Defences reduce the rate; instrumentation is what lets you improve it. Keep a golden set labelled at every stage, not only at the end: for each record, the correct fields and the correct inclusion decision. Then you can measure each stage in isolation by feeding it correct upstream input, and separately measure the end-to-end pipeline. The gap between the two tells you how much of your loss is compounding rather than any single stage being weak. Without per-stage labels you get the familiar production stall where end-to-end accuracy fell three points last week and nobody can say whether the extraction prompt, the criteria prompt, or a change in the incoming abstracts caused it. Log per-stage inputs and outputs with a correlation id from the start; retrofitting that after an incident is painful. ## Choosing where the checks go Not every seam deserves the same treatment. Put your effort where an error is both likely and unrecoverable downstream. Extraction feeding a judgement is the canonical case, because the judging stage cannot detect the error and the decision is the product. The rationale-drafting stage matters less: it is downstream of the decision and its errors are visible to a human reading the output. A useful discipline is to ask, for each boundary, what test you would write. If you can name the assertion, write it. If you cannot name any assertion for a boundary, that boundary is buying you latency and error compounding without buying control, and the two stages either belong merged or the intermediate artifact needs redesigning into something checkable. ## Recovery, not just detection When a check fails, the cheap response is to retry that stage with the failure fed back in, capped at a small number of attempts. Beyond the cap, route the record to a human queue or to a fallback path rather than letting it through. A pipeline whose only outcomes are success and silence will always look better than it is; one that can say it does not know is one you can trust the rest of the time.

  • Four stages at ninety-five percent each; what is your end-to-end accuracy and what does that assume?
    Roughly eighty-one percent, from 0.95 to the fourth power. That assumes the stage failures are independent, which is often optimistic in one direction and pessimistic in another: shared causes such as a hard or malformed source document correlate failures across stages, while a later stage that sees the original source can sometimes recover from an earlier mistake. Treat the product as a planning bound, then measure the real end-to-end number.
  • Would inserting a model-based verifier between stages be better than deterministic checks?
    Use it only where no deterministic check exists. A verifier is another fallible call that adds latency, cost, and its own error rate, and it tends to agree with plausible-looking input, which is precisely the failure you are hunting. Schema assertions and verbatim grounding checks are cheaper, faster, and do not share a failure mode with the generator. Reserve the model verifier for genuinely semantic judgements, and measure it against labels like any other stage.
  • How do you keep a chain's checks from rejecting so much that throughput collapses?
    Track the rejection rate per check as a first-class metric and set a budget for it. A check whose rejections are mostly false positives is miscalibrated and should be loosened or moved to a warning that is logged but not blocking. Route true rejections to review or a fallback rather than dropping them, so the pipeline degrades in throughput rather than in silent correctness, and revisit thresholds as prompts and models change.

saying these in an interview costs you the question

  • Assumes a later stage will notice an earlier stage's mistake
  • Treats a well-formed JSON output as evidence it is correct
  • Only measures end-to-end accuracy with no per-stage labels
  • Adds another model call to check output instead of a deterministic assertion
  • Forces the extraction stage to always produce a value with no abstain option

context