How would you judge whether rewriting a binary-protocol parser under a compile-time ownership discipline will actually close its defect tracker?
answer
- let the existing defect record decide
- sort rows by class, then count
- only validity-of-access rows close
- weight by incidents, not ticket count
- a rewrite reopens ordinary new-code bugs
basics
~20 sClassify the open defects first. The rewrite closes use-after-free, double free and dangling-reference rows by construction and nothing else; logic, retention and error-handling rows survive it. The decision is that class mix against the migration cost.
solid answer
~50 sStart from the tracker, not from the language argument. Sort the open and the recently closed rows into three buckets: validity-of-access defects, retention and leak defects, and protocol or logic defects. Only the first bucket closes by construction, and it closes completely, which is a real number you can put in front of a risk owner. Then price the other side: the patterns the old code expressed by aliasing that now need a different shape, the unchecked surface at the foreign boundary, the team's learning curve, and the fact that a whole class of silent corruption becomes a loud deterministic stop instead — better, but it is a behaviour change in production. Finally compare against cheaper moves that close part of the same bucket: scope-bound release and guard types in the existing code, a checked container wrapper, or rewriting only the frontier that parses untrusted input.
go deeper
The useful takeaway is that a safer memory model fixes one category of bug, not bugs in general. Most defect trackers are dominated by logic, not by memory validity.
Be able to sort real defects into the right buckets: which rows are about the validity of an access, which are about retention, which are about protocol logic.
Show the migration plan: convert the untrusted-input frontier first, keep the old implementation as a differential oracle on replayed traffic, and count the unchecked boundary sites up front.
Own the decision as a bet with numbers: class mix weighted by incident severity, remaining lifetime of the component, team capacity, and the alternatives that close part of the same bucket for far less.
This is a judgment call, so the answer an interviewer is listening for is a method with numbers in it, not a preference. The failure mode is arguing about the discipline in the abstract; the move is to make the existing defect record decide. ## Step one: classify the tracker Take the open rows plus everything closed in the last year or two — closed rows matter because they measure where the effort went, not just where it is stuck. Sort into buckets and count: | Bucket | Example row from a parser | Does the rewrite close it? | |---|---|---| | Validity of access | read of a frame body after its block was released | yes, by construction | | Exactly-once release | double release of a decode buffer on an error path | yes, by construction | | Interior reference | stale reference into a buffer the parser had resized | yes, by construction | | Accidental leak on an error path | release skipped on an early return | yes, through scope-bound release | | Deliberate retention | decoded frames accumulating in a session map | no | | Protocol logic | a length prefix mis-signed, a malformed frame accepted | no | | Concurrency beyond aliasing | a lock held across a slow call | no | The first four rows are the whole prize. If they are three percent of the record, the language argument has lost regardless of its merits; if they are half of it, and especially if they have caused incidents rather than tickets, the case is strong on its own. ## Step two: price the other side honestly 1. **Re-expressing what aliasing used to do.** A parser that handed out overlapping mutable views of one input buffer has to be restructured, typically into index ranges or a single writer with borrowed readers. That work is real and it lands in the hardest part of the codebase. 2. **The unchecked surface.** Wherever the parser meets a foreign boundary or externally provided memory, the guarantee is asserted by hand. Count those sites before, not after. 3. **A changed failure shape.** Where the old parser would have read past a buffer and produced a wrong answer, a bounds-checked one stops deterministically. That is usually an improvement, but it converts silent wrongness into visible unavailability, and the operational plan has to say which the service prefers. 4. **People.** The learning curve is concentrated in exactly the engineers who currently keep the parser alive. ## Step three: consider the cheaper moves that close part of the same bucket - Attach release to scope exit and wrap raw handles in guard types within the existing code; this closes most accidental error-path leaks without a rewrite. - Route all buffer access through one checked container abstraction, which closes a large share of interior-reference defects. - Build a test configuration with heavy runtime memory checking, which finds the same class of bug late instead of preventing it — cheaper, weaker. - Rewrite only the frontier that touches untrusted input, and leave the internal stages alone. The incident-weighted benefit is usually concentrated there anyway. ## Step four: state the decision as a bet A defensible answer sounds like: *these four classes are twelve of our thirty-one open rows and three of our four incidents; rewriting the frontier closes them and takes one quarter; the remaining nineteen rows are logic and retention and we keep owning them either way.* The number that makes the decision is the class mix, weighted by incident severity, and the code's remaining lifetime — a parser due for replacement in a year never earns the rewrite, no matter how the buckets fall. The pitfall to name out loud: a rewrite closes a class of defects and simultaneously opens a fresh round of ordinary new-code bugs in a component that was battle-hardened. Preserving the old parser as a differential oracle, replaying captured traffic through both, is how that cost is kept down.
- Which part of a parser would you convert first if you convert only part?The frontier that consumes untrusted input, since that is where malformed data meets buffer arithmetic and where incidents concentrate. Internal stages that operate on already-validated structures get far less benefit per unit of disruption, and the boundary between converted and unconverted code becomes one audited interface rather than many.
- What makes the rewrite clearly not worth it despite a bad memory-defect record?A short remaining lifetime for the component, or a record whose incidents trace to protocol logic rather than access validity. Also a team with no capacity to carry the learning curve on its most critical component — a half-finished migration leaves two models in one codebase, which is worse than either.
- How do you reduce the risk that the rewrite introduces new defects?Keep the old implementation running as a differential oracle: replay captured traffic and the recorded malformed frames through both, compare outputs, and only retire the old one when divergence is zero over a meaningful window. This targets the ordinary new-code bugs, which the ownership rules do nothing about.
saying these in an interview costs you the question
- Argues from language preference rather than the defect record
- Assumes the rewrite closes leak and logic rows as well
- Ignores the unchecked surface at foreign boundaries in the estimate
- Forgets that a rewrite reintroduces ordinary new-code defects
- Treats ticket counts as equal without weighting by incident severity