skip to content

How do you set a repair-attempt budget and escalation policy for structured-output failures?

level: principalimportance: should knowfreq 28%

answer

  1. measure marginal success per attempt
  2. failures are correlated, not independent
  3. the exhaustion branch matters more than the cap
  4. wrong record versus no record
  5. repair rate should trend down, not plateau

basics

~20 s

Derive the cap from measured marginal success per attempt — usually one or two, because failures are correlated and later attempts rarely convert. Then decide what happens when the budget runs out: a human queue, a partial record, or a hard failure, chosen by the cost of a wrong record versus the cost of no record.

solid answer

~50 s

Two questions decide the policy. First, **how many attempts**: measure the marginal success rate of attempt two, three and four from your own logs. It typically collapses after the first repair, because whatever survived a precise validator error is not a format slip — so a cap of one or two is the norm, and anything larger is usually paying for a schema defect. Second, and more important, **what happens at exhaustion**. That is an asymmetry judgment: for a customs declaration written to a system of record, a wrong-but-valid document is far worse than no document, so the right terminal state is a human clerk's queue with the raw output and the validator errors attached. For a low-stakes enrichment field, dropping to null and moving on is correct. Latency budgets bind independently — an interactive path may only afford one repair regardless of what the economics allow — and repair rate should be an alerted metric that trends down, not a permanent cost line.

go deeper

for a junior

Know that repair loops need a hard attempt cap and a defined fallback, so a failing request cannot retry forever or vanish silently.

for a middle

Explain why attempts are correlated rather than independent, so the second and third rarely convert, and why the failure must be logged with the raw output and validator errors whichever branch you take.

for a senior

Derive the cap from measured marginal success and set the exhaustion branch from the cost asymmetry between a wrong record and no record. Be ready to discuss retry amplification under load and per-path latency budgets.

for a principal

Own repair rate as a metric that should trend toward zero, and route the effort accordingly — schema redesign or extraction decomposition over loop tuning. Be able to argue the case where a retry is actively harmful because it manufactures a validating fabrication.

## The budget is an empirical question, not a preference The number of repair attempts is usually chosen by instinct — "we retry three times" — when it is measurable. Log, per attempt index, how many failures convert to a valid object. The shape almost everyone finds is a sharp collapse: the first repair, given a specific validator error, converts a large majority of failures, because most first-pass failures are small format slips a precise message fixes mechanically. The second converts far fewer. The third is close to noise. The reason is **correlated failure**. Attempts are not independent draws. The same model, prompt, schema and source document produce the same difficulty every time; what changes between attempts is only sampling variance and the added error message. Once the error message has been consumed and the output still fails, the residual cause is usually structural — the source genuinely lacks the fact, the schema is ambiguous, or the task exceeds the model on this input. None of those is fixed by another sample. Setting the cap where the marginal curve flattens is therefore the whole exercise, and it usually lands at one or two. ## Cost of an attempt versus cost of the outcome Each repair is a full generation: tokens in, tokens out, latency, and a slot of provider throughput. Under load that cost is not marginal — a pipeline running at a 15% failure rate with three attempts can spend a substantial fraction of its budget on retries that mostly fail. Worse, retries concentrate exactly where the system is already struggling, so a failure spike becomes a load spike, which is a classic amplification pattern. A cap is a stability control as much as a cost control. Against that sits the cost of not producing a record, which is entirely domain-dependent and is the part interviewers actually want to hear reasoned through. ## The exhaustion decision is an asymmetry judgment When the budget runs out you have four terminal states, and choosing between them is the real design question: 1. **Escalate to a human.** Right when a wrong record is materially worse than a late one. A freight forwarder's customs declaration is the archetype: a valid-looking document with a wrong HS code creates a regulatory problem, so after two failed repairs the case goes to a clerk with the source email, the model's last output and the exact validator errors attached. The clerk finishes in a minute; the alternative could cost far more. 2. **Emit a partial record.** Right when the document is mostly usable and the failing field is optional downstream. Persist what validated, mark the field unresolved with the reason, and let a later pass or a human fill it. 3. **Hard-fail the request.** Right when a caller can meaningfully handle absence — an interactive user who can retry, or a caller with a fallback path. 4. **Escalate to a different or higher-effort configuration** for one more attempt. Sometimes justified, but it is a genuinely different lever from repeating the same call, and it should be a deliberate rung with its own cost accounting rather than an automatic reflex on first failure. What unites the sound choices is that **the failure is durable and visible**: the raw output, the validator errors and the request identifier are persisted regardless of which branch you take. A repair loop that exhausts and silently drops the record is the worst outcome, because you lose both the data and the evidence. ## Latency binds independently of cost An offline batch has slack: three attempts on a small fraction of records costs minutes on a job measured in hours. An interactive request under a two-second budget may not afford even one repair, since a repair is a whole extra generation. The same failure, in the same system, can therefore warrant different policies on the batch path and the interactive path — and saying that out loud is a strong signal in an interview, because it shows the budget is derived from a constraint rather than picked. ## When a retry is worse than a failure There is a category where retrying is actively harmful. If repeated attempts nudge the model toward producing *something* that validates, a persistent budget can manufacture a plausible fabrication to satisfy a required field the source never contained. Structural validity then becomes a false signal of correctness. In those cases the honest design is a smaller budget, a nullable field, and an explicit unresolved state — the schema should permit "not found" so the model is never cornered into inventing. ## Treat repair rate as a metric that should fall The strategic framing: a repair loop is a safety net, not a component of normal operation. Instrument the rate, alert on it, and slice it by field path and document type. A field driving a disproportionate share of repairs is a schema or prompt defect whose fix removes an entire generation from every affected request forever — a far better return than tuning the loop. Teams that never look at the aggregate end up with a permanent, invisible cost line and a slowly rotting extraction quality that the loop conceals.

  • Your batch path and your interactive path share a schema. Should they share a repair budget?
    Usually not. The economics are similar but the latency constraints are not: a batch job can absorb two or three attempts on a small fraction of records, while an interactive request under a tight budget may not afford even one extra generation. The sound design shares the schema, the validator and the error formatting, and parameterizes the attempt cap and the exhaustion branch per path — interactive fails fast to a fallback, batch escalates to a queue.
  • When would you argue against having a repair loop at all?
    When the failing field is one the source may genuinely not contain, and retrying pressures the model into fabricating something that validates. Structural validity then becomes a false correctness signal. The better design is a nullable field with an explicit unresolved reason so "not found" is a legal answer, plus a much smaller budget. Also when the throughput cost of retries under load would amplify a failure spike into an outage — there, a cap of zero with clean escalation is more stable.
  • What would make you raise a cap from two attempts to three?
    Evidence, not intuition: a measured marginal conversion rate on attempt three that is high enough to beat the cost of the alternative terminal state. That is rare, and when it appears it usually means attempt two is not carrying a good error signal — the repair message is generic, or the same deterministic fix is being re-derived by the model each time. Fix that first; the extra attempt is almost always a worse buy than improving the second one.

saying these in an interview costs you the question

  • Picks a retry count by convention instead of measuring conversion
  • Treats repeated attempts as independent draws at the same success rate
  • Retries until something validates, inviting a plausible fabrication
  • Exhausts the budget and drops the record with no durable evidence
  • Uses one budget for interactive and batch paths with different deadlines

context