skip to content

An LLM pipeline decodes inputs before screening them — why does an encoded span still get through?

level: middleimportance: should knowfreq 55%

answer

  1. coverage, not closure
  2. a list somebody wrote by hand
  3. the model has no grammar to fail
  4. malformed forms pass a decoder and reach the reader
  5. which path, not which form

basics

~20 s

Decoding covers the representations somebody listed, at the stages where the step runs. A generator reconstructs forms nobody enumerated, including malformed ones a strict decoder rejects, so the step narrows the unexpanded surface rather than closing the class.

solid answer

~50 s

"We decode and then scan" is a claim about an enumeration, not about the class. A decode step expands the representations someone anticipated and wrote down, and it runs at the stages someone wired it into — in practice the user turn, rarely a field in a tool's return value or a retrieved chunk. The second reader is not a decoder at all: a generator reconstructs text statistically, has no "unsupported encoding" failure mode, tolerates nested, partial and malformed forms that a strict parser refuses outright, and its coverage grows with every capability improvement while the list only changes when somebody edits it. So what the claim buys is real — the enumerated forms genuinely die at that stage — but it is bounded by a list on one side and unbounded on the other. To someone probing the pipeline the claim is also useful reconnaissance: it tells you which forms are already dead.

go deeper

for a junior

Know that decoding first only expands the representations someone listed; anything not on that list arrives at the screening model in its original form and is scored as those characters.

for a middle

Explain the two failure modes side by side: a strict decoder refuses malformed input and passes it through, while a generator reconstructs it approximately — which is why the malformed case favours whoever placed the span.

for a senior

Move the discussion from which forms are covered to which paths the step is attached to, and note that the enumerated list and the generator's ability move on different clocks.

for a principal

Be able to say what an enumeration-based claim can honestly be asserted to provide, and why a statement that sounds like closure is really a statement about the cheapest attempts being priced out.

## What the claim actually says "We decode inputs before we screen them" sounds like closure and is really coverage. Unpacking it gives three separate limits, and each of them is a place a span survives. ### 1. A decode step is an enumeration Someone sat down and listed the representations worth expanding, then wired in a routine for each. That list is finite and static. It grows when a person edits it, usually after an incident. Everything not on it passes through the step unchanged and reaches the screening model in its original form — where it scores as whatever those characters look like. The second reader has no list. A generator reconstructs a representation because reconstruction is the most plausible reading of the input, not because a branch matched. Its coverage is a smear rather than a set: nested forms, forms it has seen a handful of times, one representation carried inside another, a form applied to only part of a span. You cannot read its coverage off a specification, and it is not the same on two model releases. ### 2. A strict decoder and a statistical reader fail differently This is the sharpest part of the answer. A real decoder has a grammar. Handed input that does not conform — wrong padding, a stray character, a truncated tail — it raises an error or refuses, and the pipeline's usual behaviour is then to pass the value through untouched, because refusing the request outright would break ordinary traffic. A generator has no such failure mode. It will reconstruct a mangled form approximately, and "approximately" is often enough, because an instruction only has to be legible, not byte-exact. So the deliberately malformed case is worse than the clean one: it is the case where the decode step is guaranteed not to expand the span and the model quite likely will. This is *not* the classic canonicalise-before-validate ordering bug, and confusing the two is the common wrong turn. In that bug both sides are explicit decoders with grammars, and the defect is that they disagree about the same input. Here there is no second decoder at all — the second side is a model, and reasoning about it from a specification does not work. ### 3. The step runs somewhere, and "somewhere" is usually the user turn A context window is assembled from many sources. The user's own message is the one a team labels untrusted, so that is where the input handling gets attached. The rest arrives on other paths: a retrieved chunk, a fetched page, another component's output, a field in the JSON a backend hands back to a tool call — a note written long ago by an end user of some host product, treated by everybody as system data. For a feature a platform vendor embeds inside other companies' products, that material is also invisible to the vendor by construction: the vendor sees the prompt it assembled, not what the host product's data flow put into it. A span placed on such a path never meets the decode step at all. Its author does not need an exotic representation — only one the screening model does not read as a directive. ## The boundary moves, and it moves the wrong way The honest version of "we decode and then scan" is: *we expand the representations we anticipated; the generator expands the ones we did not.* And the two sides move on different clocks. The list changes on release cycles driven by findings. The generator's reconstruction ability improves with every capability jump, and the screening model — smaller, cheaper, run on every request under a latency budget — improves more slowly. Each capability release therefore tends to widen the gap the technique lives in rather than narrow it. ## What this looks like from the other chair Someone probing the pipeline reads the claim as free reconnaissance. Published coverage is a list of forms that are already dead; the interesting region is what borders it. The cheap probes are not exotic representations at all — they are the ones that sit at the edges of a grammar: partially applied, nested one level, malformed in a way that a parser rejects and a reader shrugs at. And the highest-value question is not *which form* but *which path*, because a path that the step was never attached to makes the whole enumeration irrelevant. ## What the claim is still worth It is not nothing. Every representation on the list genuinely stops working at that stage, which removes the cheapest and most-copied forms and raises the cost of a first attempt. The mistake is only in the word "then": the sentence describes a stage that expands a known set, and it is heard as a sentence about the class of all encodings.

  • Why is this not the same as a parser differential between two decoders?
    In a parser differential both sides are explicit decoders with grammars and the defect is that they disagree about one input. Here the second side has no decoder: a generator reconstructs statistically, has no unsupported-encoding failure, and tolerates nested, partial or malformed forms a strict parser refuses. You cannot derive its coverage from a specification, which is what makes the comparison misleading rather than merely imprecise.
  • Where do teams usually attach the decode step, and why does that matter more than the list?
    To the user turn, because that is the text they label untrusted. Everything else — a retrieved chunk, a fetched page, another component's output, a field in a tool's return value — reaches the context window on paths that never run it. A span placed on one of those paths makes the enumeration irrelevant, so path coverage is the stronger question and it is the one teams answer least often.
  • Does a deliberately malformed form help or hurt whoever placed the span?
    It usually helps, which is the counterintuitive part. A strict decoder refuses to expand input that does not conform and the pipeline passes the value through untouched, while a generator reconstructs a mangled form approximately — and approximate is enough, because an instruction only needs to be legible, not byte-exact. The malformed case is the one where the step is guaranteed to do nothing.

saying these in an interview costs you the question

  • Treats a decode step as closing the class rather than a list
  • Calls it the canonicalise-before-validate bug with two decoders
  • Assumes the decode step runs on every path into the context window
  • Thinks a malformed encoding is harder for a model than a clean one
  • Assumes the decoder list and the model's ability move together

context