You inherit a large system that moves natively serialized objects between services, into a shared cache, and into client-held session blobs. How would you eliminate this bug class rather than patch instances, and what holds the line while the migration runs?
answer
- inventory decode sites, not endpoints
- delete the requirement first
- version tag → dual read → drain → delete reader
- lint ratchet as the burndown metric
- signing ≠ safety; budgets always
basics
~20 sInventory decode sites and their producers, set a target of schema-first codecs binding into code-chosen types, migrate behind a version tag, and delete the requirement where possible — client blobs become server-side references. Meanwhile: allow-list filters, budgets, least-privileged decoders, banned-API lint.
solid answer
~60 s**Target state, stated first:** no component decodes a self-describing encoding from a producer it does not fully trust. Payloads are plain data bound into a type the code named; where several shapes exist, a finite code-owned tag→type table supplies them. **Get there in four moves.** 1. **Inventory decode sites**, not endpoints — RPC codecs, cache and session stores, queues, config and plugin loaders, uploads — and record producer trust, storage hop, decoding-process privilege and current defence rung. 2. **Delete requirements before defending them.** A client-held blob becomes an opaque id resolved server-side; that removes the boundary rather than hardening it, and it is usually the single biggest reduction. 3. **Migrate without a flag day**: version-tag the envelope, dual-read both codecs, cut writers over, then delete the old reader. Deleting the old reader is the step that ends the class; a permanent dual-read is a permanent vulnerability. 4. **Hold the line meanwhile**: platform type filters configured as allow-lists, resource budgets, decode in least-privileged egress-denied processes, and CI lint banning the legacy decode APIs so the count cannot grow. **Exit criterion**: decode sites accepting non-fully-trusted producers, tracked to zero.
go deeper
Understand the target state — plain data bound into a type the code chose — and that the fix is a migration, not a patch.
Describe the version-tag, dual-read, drain, delete-reader sequence and know that budgets and allow-list filters are the interim controls.
Own the inventory and the ranking, run the migration across storage hops, and insist that the legacy reader deletion is the completion criterion.
Argue the whole programme: eliminate requirements first, make the safe codec the platform default, ratchet with CI lint and an exception register, and report a single decreasing metric rather than closed findings.
## Frame it as class elimination, not finding closure The treadmill answer — upgrade libraries, add the newest blocklist, rerun the scanner — loses by construction, because the space of reconstruction-time behaviour in your dependency graph is open-world and grows with every upgrade. The winning frame is: **change the mechanism so the sender can no longer choose the type**, then measure the shrinking population of sites that still can. ## Step 1 — inventory decode sites Endpoints are the wrong unit. Enumerate every place bytes become objects: - inter-service RPC codecs and message-queue payloads - cache entries, especially shared caches written by one service and read by another - session and saved-state stores, and anything the client holds - database columns holding serialized graphs - configuration, feature-flag and plugin loaders - file uploads, import jobs, backup and restore paths For each record: who can influence the bytes (including second-order producers), whether a storage hop separates producer from consumer, the privilege of the decoding process, the loaded type surface, and which rung of the defence ladder currently applies. This table *is* the programme plan; it also gives you the metric. ## Step 2 — delete requirements before defending them The cheapest site to secure is one that stops existing. - **Client-held state** is the clearest case. Replace the blob with an opaque, unguessable identifier resolved against server-side storage. You lose statelessness at the edge and gain the removal of an entire trust boundary, plus straightforward revocation. - **Objects between services** were usually a convenience, not a requirement. Most such interfaces carry a handful of fields; a schema-first contract is smaller, versionable and faster. - **Shared caches used as an integration channel** should be narrowed to one writer and one reader, or replaced with an explicit API — the value is as much about coupling as about security. ## Step 3 — migrate without a flag day The mechanics that survive contact with a live estate: 1. Wrap payloads in an envelope with an explicit codec version tag. 2. Deploy readers that accept both codecs, selected by the tag — never by sniffing bytes, which re-introduces guessing. 3. Move writers over, one producer at a time, with the ability to roll back to the old codec. 4. Drain the storage hops: caches must expire, queues must empty, stored columns must be rewritten by a backfill. This is the step that gets forgotten and the reason a "finished" migration still has a live decoder. 5. **Delete the legacy reader** and remove the dependency. Until then, nothing has been eliminated — a dual-read service is fully exposed via the legacy branch. In a system with no external consumers to protect, prefer a straight cutover over a compatibility window: every shim is code a future reader has to reason about, and a legacy reader left "just in case" is the vulnerability. ## Step 4 — hold the line while it runs These are interim controls, and calling them interim out loud is part of the answer: - **Platform type filters configured as allow-lists** for each legacy decode site. Configured as deny-lists they are telemetry only. - **Resource budgets** everywhere at once — size, decompressed size and ratio, nesting depth, element counts, decode timeout, concurrency cap. Cheap, uniform, and they cover the exhaustion band that type control does not. - **Least-privileged decoding**: run legacy decoders in processes with narrow credentials and denied egress, so a successful chain has nowhere to go. This caps blast radius; it does not fix the defect, in the same way a least-privileged database account caps injection without preventing it. - **A ratchet**: CI lint that bans the legacy decode APIs outside an explicit, reviewed exception list. The count of exceptions is your burndown chart and it must be monotonically decreasing. - **Detection last**: a canary type in each allow-list that nothing legitimate uses, alerting when instantiated. ## Controls people propose that do not close the class - **Signing the blob.** It answers who produced the bytes, not whether decoding them is safe. It has no freshness, so old blobs replay; it makes every signer and its build pipeline part of your trusted computing base for code execution; and where the key must ship to clients it is worthless. Useful as a boundary control on top of a rung, never as the plan. - **Encrypting the blob.** Confidentiality says nothing about the plaintext's origin or safety. - **A web application firewall.** Byte-level pattern matching against nested, compressed, re-encoded structured data is evadable and, worse, it produces a green dashboard. ## How you would report progress One number: decode sites that accept input from producers you do not fully trust, broken down by defence rung, trending to zero — with a second line for sites still on rung 3 and the date their legacy reader is deleted. Anything phrased as "vulnerabilities fixed this quarter" is measuring the treadmill.
- The team wants to keep the legacy reader permanently "for compatibility". What is your answer?A dual-read service is fully exposed through its legacy branch, so the migration has produced cost without removing risk. Compatibility windows are for consumers you cannot redeploy; inside one estate you can, so pick a date, drain the storage hops, and delete the reader. If some producer genuinely cannot be moved, isolate it: a dedicated least-privileged decoder process with an allow-list and budgets, tracked as an explicit exception with an owner and an expiry.
- How do you keep the class from coming back after the migration?Make the unsafe path harder to reach than the safe one. Ship a platform codec that is safe by default, ban the legacy decode APIs in CI with a reviewed exception list, put the check in the service template so new services inherit it, and add the question to design review: does any component decode a self-describing encoding from a producer it does not fully trust? Detection alone will not hold — the constraint has to live in the toolchain.
- What would you do first if the estate is too large to migrate soon?Two things in parallel: apply resource budgets and allow-list-configured type filters everywhere, since they are uniform and cheap, and move only the top of the ranked list — sites reachable by untrusted producers whose decoding process is privileged. Meanwhile, install the lint ratchet immediately so the population stops growing while you work through it.
saying these in an interview costs you the question
- Planning around library upgrades and refreshed blocklists as the remediation.
- Declaring a migration done while a legacy reader is still deployed or blobs remain in caches, queues and columns.
- Proposing signing or encryption of the blob as the strategic fix.
- Relying on a network filter to match patterns inside nested or compressed structured payloads.
- Measuring progress in findings closed rather than in decode sites remaining.