Rank the defences available when a service must accept serialized data from an untrusted producer, from strongest to weakest, and justify at least one adjacent pair. Where does a blocklist of known-dangerous types sit, and why does it sit there?
answer
- separation > enumeration > validation > detection
- closed-world beats open-world
- allow-list tag → code-owned type map
- budgets are orthogonal to types
- signing = origin, not safety
basics
~20 sStructural separation first: a format with no type channel, bound into a code-chosen type. Where polymorphism is unavoidable, closed-world enumeration takes the top rung. Then validation — blocklists, size and depth caps — which is open-world heuristic. Detection last.
solid answer
~60 sThe ladder, strongest to weakest: 1. **Structural separation** — decode with a parser that has no type or behaviour channel, then bind the resulting inert tree into a type the code named. The attacker cannot select a type because there is no channel to select one through. This is a guarantee, not a filter. 2. **Closed-world enumeration**, which occupies the top rung when structural separation is genuinely unavailable (a legacy codec, real polymorphism on the wire). A finite, code-owned map from tag to permitted type. Finite and stable under dependency churn. 3. **Validation** — type-name blocklists, patterns, size/depth/element budgets. Open-world: it constrains what you enumerated while the admissible space stays unbounded. Resource budgets belong here and are still mandatory, because exhaustion needs no gadget. 4. **Detection** — canary types in the allow-list, alerting on rejected decodes, egress control on the decoding process. It catches attempts; it prevents nothing. Enumeration beats a blocklist for one reason: closed-world versus open-world. A blocklist still admits every dangerous type you didn't name, including ones the next dependency upgrade introduces.
code
text · 9 linestree = parse_plain_data(bytes) # no type channel, inert output
enforce_budgets(tree) # depth, element count, size, time
PERMITTED = { "order.created": OrderCreated,
"order.cancelled": OrderCancelled } # finite, code-owned
type = PERMITTED[tree["kind"]] or reject(400) # closed world
msg = bind(tree["payload"], type) # code chose the type
validate_business_rules(msg) # plain values nowgo deeper
Know the top and bottom: prefer a plain-data format bound to a type the code chose; understand that a list of banned classes is not a fix.
Reproduce the ladder in order and give the open-world versus closed-world reason for enumeration beating a blocklist; remember budgets are separate from type control.
Say which rung each sink in your system is on, distinguish allow-list from deny-list configuration of the same platform filter, and treat signing as a boundary control rather than a rung.
Set the policy: which components may accept self-describing encodings at all, what the enumerated set is, who owns it, how it is reviewed, and what the exit path is for legacy codecs.
## Why rank at all Defences are not interchangeable mitigations to be stacked in any order. They weaken from **guarantee to heuristic** as you go down the ladder, and knowing which rung you are on tells you what you may claim in a design review. The same ordering used across this tier applies here: structural separation > escaping/transformation > validation > detection — with the important adaptation that deserialization has no meaningful "escaping" rung (there is no lexer to escape into), so **closed-world enumeration** occupies the rung below structural separation. ## Rung 1 — structural separation Remove the channel. Use a decoder whose grammar can express only values: strings, numbers, booleans, null, lists, maps. The output is inert. Then bind that tree to a type your code named at the call site. The attacker cannot choose a type because the format has no way to say one, and reconstruction runs only the code you wrote. This is a guarantee in the same sense that a compiled query template with bound values is a guarantee: it is a property of the mechanism, not of a rule you maintain. It survives dependency upgrades, new gadget research, and encodings you did not anticipate. A schema-first codec (an interface definition compiled into types on both sides) is the same rung reached by a different road: the wire carries field numbers or names, never type identity. ## Rung 2 — closed-world enumeration Sometimes the wire genuinely must carry a choice: a message envelope with several payload shapes, or a legacy protocol you cannot yet replace. Then the document supplies a short **tag**, and your code holds a finite table mapping tag to permitted type. The document *selects from your list*; it does not *name a type*. The reason enumeration outranks anything below it is open-world versus closed-world. A blocklist, a regex on the type name, or a package-prefix rule constrains characters or names — while still admitting an unbounded set, which includes every class in every library you ship now or later. An enumerated table is finite, readable, reviewable and owned by the code, so its risk changes only when a human edits it. Platform-level type filters that accept an allow-list belong on this rung. The same filter configured as a deny-list drops to rung 3, which is the trap: the same API can express either, and only one of them is a control. ## Rung 3 — validation Everything heuristic: blocklists of known-dangerous types, name patterns, and the resource budgets — maximum encoded size, maximum nesting depth, maximum declared collection or array size, maximum element count, decode timeout, and a cap on decompressed size for compressed payloads. Two honest things must be said about this rung. First, blocklists are unsound; use them only as defence-in-depth or as a temporary tourniquet, never as the argument that a sink is safe. Second, **resource budgets are mandatory even at rung 1**, because a decoder that constructs only permitted types can still be told to construct billions of them, to recurse a thousand levels deep, to expand a compression bomb, or to fill a hash container with colliding keys. Type restriction and resource restriction are orthogonal. ## Rung 4 — detection Count and alert on rejected decodes; plant a canary type in the allow-list that nothing legitimate uses and page when it is instantiated; deny outbound network from the decoding process so a successful chain has nowhere to go; run decoding in a least-privileged process. None of this prevents the first exploitation; all of it shortens the window and raises the cost. ## The controls that are not rungs - **Signing or MAC-ing the blob** moves the trust boundary; it does not change what the decoder does. It buys authenticity of origin, not safety of content. It is worth little when the key ships to clients, when any low-privilege component can sign, or when there is no freshness (a replayed old blob verifies perfectly). And a single key compromise converts directly into code execution, because the verified blob is then fully trusted. Treat it as a boundary control layered on top of a rung, never as the rung. - **Encrypting the blob** is weaker still for this purpose: confidentiality says nothing about who produced the plaintext. - **Least privilege on the process** caps blast radius; it does not stop the defect, exactly as least-privileged database accounts cap injection without preventing it. ## Saying it well Name the rung you are on and what it guarantees. "We are on rung 1 for the public API — plain-data parse and bind to a named type — and on rung 2 for the internal queue, where the envelope carries a tag resolved through a nine-entry table. Budgets are enforced on both. The old blocklist stays as telemetry, and it is not part of the safety argument."
- Your platform offers a global deserialization filter. Does configuring it make the sink safe?Only if it is configured as an allow-list, and only for the types it actually covers. Configured as a deny-list it is rung 3 — useful telemetry and defence in depth, not a safety argument. Also check its scope: a process-wide filter may not apply to a nested or delegated decoder, and it typically does nothing about resource exhaustion, so you still need budgets.
- The producer signs every blob with a key only it holds. Can we skip the type allow-list?No. A signature answers "who produced these bytes", not "is decoding them safe". You have made the producer, its key material, its build pipeline and anyone who can compromise it part of your trusted computing base for code execution. Add that a signature carries no freshness, so an old blob replays cleanly. Keep the ladder rung and treat the signature as a boundary control on top.
- Which failure remains after a perfect type allow-list?Denial of service and object-state tampering. Exhaustion needs no gadget — depth, declared collection sizes, decompression, hash flooding. And if a permitted type carries authorisation-relevant state such as a role or a price, a tampered payload changes behaviour without executing anything, so decoded objects must still be treated as untrusted input and re-authorised.
saying these in an interview costs you the question
- Presenting a blocklist as the fix rather than as telemetry.
- "We upgraded the library, so the gadgets are gone" — a treadmill against an open world.
- Assuming a type allow-list stops resource-exhaustion attacks.
- "It's signed, so we can deserialize it" — conflating authenticity with safety.
- Enforcing budgets only on encoded size, ignoring nesting depth, declared collection sizes and decompression ratio.