skip to content

Rewriting the proprietary parts of a system is estimated in engineer-months — why is that number systematically low?

level: seniorimportance: should knowfreq 44%

answer

  1. the interface is the cheap part
  2. you inherited behaviour, not just an API
  3. the estimate counts calls, not consequences
  4. unknowns sit on the platform nobody has operated
  5. size it by porting one component first

basics

~20 s

The estimate counts call sites, while the work is rebuilding what the component did behind them — retries, ordering, durability, scaling and the operational tooling around it. Those were the reasons it was adopted, and they are discovered during the port, not before.

solid answer

~40 s

A proprietary component was adopted because it removed work, so replacing it means putting that work back. Swapping the calls is a day; re-creating the delivery guarantee, the scaling behaviour, the failure handling and the operational surface underneath them is the project. Three things then push the number down further: the estimate is produced by people who know the platform being left and not the one being joined, so the unknowns all sit on the side nobody has operated; the parts with no equivalent are found while porting rather than while planning; and proving the replacement behaves the same under production load is a second body of work that rarely appears in the first figure. Size it by porting one representative component end to end, then scale — not by counting call sites.

go deeper

for a junior

Remember that replacing a managed component means rebuilding what it did for you — retries, ordering, scaling — and not merely changing the lines that called it.

for a middle

Explain which guarantees a component supplied and what has to appear in your own code when the replacement supplies a different one.

for a senior

Show the method: classify dependencies by whether an equivalent exists, port one end to end to get a real multiplier, and budget the equivalence proof as its own line.

for a principal

Speak to the bias rather than the number. Say who produced the estimate, which side of the move their expertise is on, and what evidence would let you widen or narrow the range.

## The interface is the cheap part When a system is built on a platform's proprietary component — a managed queue, a managed store, a managed identity handoff, a managed scheduler — the code that touches it is small. A dependency search finds every call site in an afternoon, and that number feels like the estimate. It is not, because the calls were never the value. **What you bought was everything you did not have to write behind them**, and that is exactly what a rewrite has to restore. Put it in the acquisition case: a media transcoding service is being moved off a platform so two estates can be merged. The pipeline hands work off through a managed queue with an at-least-once delivery guarantee, per-key ordering, a dead-letter path and automatic scaling of consumers. The code that enqueues and dequeues is perhaps two hundred lines. The behaviour behind those two hundred lines is the part the target platform may deliver differently, or not at all. ## Where the missing engineer-months live - **Semantics you inherited.** Delivery guarantee, ordering, visibility and retry behaviour, durability, consistency on read. When the replacement offers a different guarantee, the difference does not disappear; it moves into your code as deduplication, sequencing or compensation. - **Behaviour you never implemented.** Scaling with load, backpressure, poison-message handling, throttling and backoff. These were operational defaults, so nobody wrote them down and nobody counted them. - **The operational surface.** Dashboards, alerts, the runbook, the audit record, the on-call knowledge of what normal looks like. The replacement starts with none of it. - **Data shape and semantics.** Exported data lands in a shape the replacement reads differently — identifier semantics, ordering, time resolution, encoding of absence. Reconciling that is its own workstream. - **Proof of equivalence.** A rewrite is not finished when it compiles. Running both paths on shadow traffic and comparing outputs is usually the longest single phase, and usually absent from the first estimate. ## Why the bias points one way Estimates are usually produced by the team that operates the current platform. Their expertise is on the side being left, so every unknown is on the side being joined; and unknowns resolve upward far more often than downward. Meanwhile the parts with no equivalent are discovered while porting, because they are precisely the parts nobody thought about — they worked. | What the estimate counts | What the work turns out to be | |---|---| | Call sites to a proprietary interface | The guarantees and failure behaviour behind that interface | | Code that must compile against something new | Code that must behave the same under production load | | Components with a like-for-like replacement | The small minority with none, which dominate the effort | | Nominal engineer-months of the team | Those months minus the roadmap the team still owns | The last row is worth saying aloud in a review. A team of six does not supply six engineer-months a month to a migration; it supplies whatever is left after the product work that did not stop. ## What a defensible estimate looks like 1. **Classify, do not count.** Sort every dependency into: moves unchanged, has a close equivalent, has no equivalent. Effort concentrates almost entirely in the third bucket, and its size has nothing to do with line count. 2. **Port one representative component end to end** — code, data, tests, comparison run, alerting — and measure what it actually took. That measurement is the only honest multiplier you will get. 3. **Quote a range with its assumptions attached**, not a point. A point estimate for work nobody on the team has done before is a guess with a decimal place. 4. **Re-estimate after the first component lands**, and treat a large correction as information rather than as failure. The first port exists to buy the estimate. 5. **Budget the equivalence proof explicitly**, as its own line, so it is not the thing quietly dropped when the date gets tight. One more consequence for the wider exit number: because the rewrite is the line most likely to slip, and because slip lands on the overlap where two platforms are billed together, an optimistic rewrite estimate is not a local error. It is the input that makes the whole exit estimate wrong, through a line it does not appear in.

  • How would you produce a rewrite estimate you would defend to a budget owner?
    Port one representative component end to end on the target, with its tests, alerting and a comparison run against the current one. Use the measured effort to scale the rest, weighting each remaining component by how much platform behaviour it leans on rather than by size, and quote a range with the assumptions beside it. Then re-estimate once the first port has landed.
  • Most of the code usually moves untouched — why does that not shrink the estimate much?
    Because effort does not follow line count. Code speaking widely implemented protocols and formats often is the bulk of the system and moves nearly as it is, while the work concentrates in the small minority that leaned on platform behaviour. Keeping ninety percent of the code can still leave most of the engineer-months in front of you.

saying these in an interview costs you the question

  • Estimates the rewrite by counting call sites to the proprietary interface.
  • Assumes a like-for-like replacement exists on the target for everything.
  • Quotes a point estimate for work nobody on the team has done before.
  • Leaves out proving the replacement behaves the same under production load.
  • Counts nominal engineer-months without subtracting the roadmap the team still owns.