skip to content

Why does an LLM's long structured output drift, like row 60 of a 100-row table losing a field?

level: seniorimportance: should knowfreq 44%

answer

  1. adherence decays with distance
  2. the model copies its own last rows
  3. one slip propagates forward
  4. valid JSON, missing column
  5. shorten the generation, assert the count

basics

~20 s

Adherence decays over a long generation: the format spec recedes into distant context, each new row is copied from nearby rows rather than the schema, and one omission propagates. Fix by batching, per-row validation with row-count postconditions, and constrained decoding.

solid answer

~50 s

Drift is a decay problem, not a syntax problem. The schema instruction sits thousands of tokens back by the time row sixty is generated, and the strongest local signal is the rows immediately preceding it — so the model self-conditions on its own recent output. Once one row omits a field, the following rows copy that shape, and a fifty-row tail quietly loses a column while the response still parses as perfectly valid JSON. Truncation at the output-token limit looks similar but is a different bug. Mitigations stack: split the job into batches of ten or twenty rows so no generation runs long, repeat the schema at the head of each batch, and constrain decoding per row so a required key cannot be skipped. Then enforce postconditions in code — expected row count, every required key present in every row — because a parse check alone will not catch it.

code

python · 5 lines
python
def check_batch(rows: list[dict], expected: int, required: set[str]) -> None:
    assert len(rows) == expected, f"expected {expected} rows, got {len(rows)}"
    for i, row in enumerate(rows):
        missing = required - row.keys()
        assert not missing, f"row {i} missing {sorted(missing)}"

go deeper

for a junior

Understand that a long list from a model can quietly change shape partway through, and that the response parsing successfully does not mean every row is complete. Check row counts.

for a middle

Explain the mechanism: the format instruction is far back in context while the previous rows are close, so the model copies its own recent output and one omission propagates forward through the rest of the generation.

for a senior

Demonstrate the operational fix — batch the generation, constrain per row, assert row counts and required keys in code, repair only the failing rows, and track per-column null rates by position to catch semantic drift that passes validation.

for a principal

Own the cost curve. Batching multiplies prompt tokens and request count, so argue where prompt caching, item-level constraints and postcondition sampling sit relative to the accuracy the pipeline actually needs, and set the data-quality SLO the drift metrics report against.

## The symptom You ask for a hundred rows of structured extraction. The response parses cleanly. Your schema validator, if it only checks the top-level array, passes. Days later someone notices that the last forty entries have no `allergens` field, or that a column of prices silently became strings, or that after row sixty the model started abbreviating category names it had been spelling out. Nothing errored. That is format drift: the shape degrades gradually across a long generation, and every automated check that looks at the response as a whole is blind to it. ## Why it happens Three mechanisms combine. **The instruction recedes.** Autoregressive generation conditions on everything before the current token, but influence is not uniform. By row sixty, the format specification is thousands of tokens back, competing with the model's own output for attention. The nearest, most repetitive, most recent evidence for "what a row looks like" is the previous rows — so the model copies them. **Self-conditioning compounds a single slip.** This is the important consequence. One omitted key is a local sampling accident. But that row is now in the context, and the next row is generated in its image. Errors are absorbing rather than self-correcting: drift is monotonic in practice, which is why the failure looks like "everything after row sixty" rather than scattered bad rows. **Long repetitive generation degrades on its own.** Producing hundreds of near-identical rows is a regime where models become terser, start abbreviating, or begin skipping fields whose values repeat — a compression instinct that a human note-taker would share. Distinguish this from **truncation**: hitting the output-token cap ends the response mid-object and usually produces a parse error, not a subtly shortened row. Check the finish reason before diagnosing drift. ## What constrained decoding does and does not fix A grammar or schema constraint at the *item* level prevents structural drift: if the row grammar requires all fourteen keys, a row cannot be emitted with thirteen. That closes the dropped-field class entirely, and it is the strongest single mitigation available. It does not close semantic drift — the model can satisfy the schema by writing `null`, `""` or `"unknown"` in a field it has stopped bothering to extract, and by row eighty the null rate on a column can climb without any structural violation. So constrained decoding moves the problem from "my parser explodes" to "my data quality quietly degrades", which is better but still requires measurement. ## Mitigations, in order of leverage **Do not generate long structured output in one call.** Batch the work: ten or twenty rows per request, schema repeated at the head of each batch, results concatenated by your code. This bounds how far any generation can drift and makes each unit independently retryable. The cost is repeated prompt tokens, which prompt caching of the shared prefix largely absorbs. **Constrain per item, not per document.** Applying the grammar to each row makes required keys non-optional at the sampler level. **Enforce postconditions in code.** Assert the expected row count, assert every required key is present in every row, and assert type consistency down each column. Also track per-column null and empty rates by row position — a rising null rate in the back half is the fingerprint of semantic drift and no per-row validator will flag it. **Repair narrowly.** When a batch fails a postcondition, re-request only the failing rows, quoting the specific defect ("rows 12-20 omitted `price_cents`; return those nine rows only, all fourteen keys present"). Re-running the whole hundred-row job is slower, more expensive and just as likely to drift somewhere else. **Give each row an identity.** Requiring a stable key — a source line number or an input-supplied id — lets you detect not just dropped fields but dropped, duplicated and reordered rows, which are the sibling failures nobody instruments until they get burned. ## The judgment an interviewer is listening for The weak answer is "tell the model to be consistent" or "lower the temperature". Temperature zero does not remove drift; it removes variety, and a deterministic decode can walk into the same degradation every time. The strong answer recognizes that a long generation is a compounding process, that you shorten it rather than exhort it, and that the only trustworthy detector is a postcondition in your own code — the response parsing successfully proves nothing about rows fifty through a hundred.

  • How would you tell drift apart from the model simply hitting the output-token limit?
    Check the finish reason on the response. A length stop means truncation, and the tell is a response that ends mid-token or mid-object and fails to parse. Drift produces a complete, well-formed response whose later items are structurally or semantically thinner. Logging finish reason alongside row count separates the two without guesswork, and they need different fixes — a bigger cap or batching versus per-item constraints.
  • Constrained decoding guarantees every key is present. What drift can still get through?
    Semantic drift. The model satisfies the schema with null, an empty string, a repeated value copied from the previous row, or a progressively abbreviated label. Track per-column null and empty rates against row position: a flat rate in the first quartile and a climbing one in the last is the signature. Per-row schema validation will never flag it because nothing is invalid.
  • Why is re-requesting only the failing rows better than retrying the whole batch?
    It is cheaper, faster, and it shortens the generation — which is the actual cause. A narrow repair also lets you quote the specific defect, which is far stronger conditioning than restating the schema. Retrying a hundred rows re-runs a long generation that is just as likely to drift somewhere else, and you lose the rows that were already correct.

saying these in an interview costs you the question

  • Assuming valid JSON means every row is complete
  • Claiming temperature zero eliminates format drift
  • Generating hundreds of structured rows in a single call
  • Confusing output-token truncation with gradual drift
  • Validating only the top-level array, never each row

context