Before a team rewrites a working single-machine data job, which two costs are routinely left out of the comparison?
answer
- price the hour before the rewrite
- engineering days, all four parts
- the peak step sets the size
- two implementations, changed twice, forever
- a retirement date bounds the cost
basics
~20 sTwo: the hourly price of renting a machine several sizes larger, weighed against the engineering days the rewrite genuinely takes; and the standing cost of two implementations of the same logic, which must be changed twice and reconciled for as long as both exist.
solid answer
~50 sThe first missing number is the cheap one. A machine several sizes larger — more memory and cores on one box, rented by the hour — has a price that can be looked up in minutes, and for a job that runs once a night the relevant figure is hours per month, not a monthly rate. Against it sits the rewrite's real cost in engineering days, which is almost always underestimated because people price writing the new code and not validating it against the old. The second missing number is permanent: once a second implementation of the same logic exists, every change to that logic is made twice and reconciled, and any divergence between the two outputs has to be adjudicated by someone. That cost does not appear in the project plan because it starts after the project ends.
go deeper
Recall that renting a bigger machine is an option with a real price, and that having two programs computing the same number is not free even when both of them work correctly.
Explain why the peak at one step, rather than the size of the dataset, determines how much larger a machine needs to be, and why an hourly price for a nightly job is a small number.
Show that you price the whole rewrite — validation, surrounding operational work and the learning curve, not just the code — and that you expect reconciliation between two implementations to be recurring work.
Own the part nobody funds: whether a second implementation is a bounded migration with a retirement date and an owner, or an unbounded standing cost adopted by accident, and say which one the plan is actually proposing.
## Two numbers that never reach the slide When a data job outgrows the way it is currently written, the comparison that gets made is usually between *the job as it is today* and *the job as we imagine it after the rewrite*. Two costs are missing from that comparison almost every time, and they push in opposite directions. The first is the price of simply running the same program on **a machine several sizes larger** — renting more memory and more cores on one box, priced per hour — set against the engineering days a rewrite actually consumes. The second is the **standing cost of two implementations of the same logic**, which begins the day the second one works and never ends while both exist. ## Pricing the larger machine honestly The reason this number is missing is rarely that it is hard; it is that nobody looks it up. When you do, three things usually change the picture: - **Rent is charged for the time you use it.** A job that runs for ninety minutes each night uses roughly forty-five hours a month. The relevant figure is the hourly price times the hours actually consumed, not a headline monthly rate for a machine left running. - **The step that failed sets the size, not the dataset.** The measurement that names the peak at a particular step tells you how much headroom is needed. Sizing to the whole dataset when one step's peak was the problem buys a much bigger machine than the situation requires. - **There is a ceiling, and it is worth knowing where it is.** Rentable sizes stop somewhere. A larger machine that buys eighteen months of room is a very different proposition from one that buys six weeks, and which of those it is can be worked out from what the measurement says rather than guessed. None of this makes the larger machine the answer. It makes it a **priced option** rather than an unpriced one, which is the only state in which a comparison means anything. ## What a rewrite actually costs The symmetric failure is to price a rewrite as the time to write the new code. The honest figure includes: 1. **Writing it** — usually the only part estimated, and usually the smaller part. 2. **Proving it produces the same answers** — running both over the same inputs and reconciling the outputs, including the rows where they legitimately differ and the rows where the original was quietly wrong. 3. **Rebuilding the surroundings** — the scheduling, the alerting, the failure handling and the operational habits that the original accumulated over years and that nobody wrote down. 4. **The learning curve** — the first months during which the team is slower at the new thing than it was at the old thing. The dangerous property of these four is that only the first is visible when the plan is written. ## The standing cost of two implementations This is the cost that is genuinely permanent, and it is worth stating plainly: **from the moment a second implementation of the same logic exists, that logic has to be changed twice and the two have to be made to agree.** | what happens | with one implementation | with two | |---|---|---| | A rule changes | edited once, deployed once | edited twice, reviewed twice, deployed twice, and the two edits must mean the same thing | | The outputs disagree | cannot happen | somebody must decide which one is right, and that person needs both in their head | | A new person joins | learns one program | learns two, plus the differences that are deliberate and the ones that are accidents | | One is retired | not applicable | requires a decision and a dated plan, which is exactly what tends not to happen | The reconciliation is the expensive part, because it is not a task with an end. Two programs computing the same number from the same input diverge for boring reasons — a rounding difference, a different disposition for absent values, one of them dropping a record the other kept — and each divergence costs a person a day to adjudicate and often produces no change to either program. ## What varies Not every second implementation carries the full standing cost. Where it is created **deliberately and temporarily**, with a written date for retiring the original and someone accountable for that date, it is a migration and the cost is bounded. Where it is created as a rewrite that 'we'll switch over once we trust it', the original tends to survive indefinitely as an unfunded fallback, and the standing cost becomes permanent by default. The difference between those two outcomes is a retirement date and an owner, not a technical property. ## What to write down Before either conversation begins, three things should be on paper: the hourly price of the larger machine multiplied by the hours the job genuinely runs; the engineering days for all four components of the rewrite, not just the first; and an explicit statement of whether a second implementation will exist temporarily, with a date, or indefinitely. Any comparison made without those three is a comparison between a number and a feeling.
- The job runs for ninety minutes once a night. How does that change the price of a larger machine?Substantially, because rent is charged for time used. Forty-five hours a month at an hourly price is a very different figure from a machine left running continuously, and it is often small next to a single week of engineering time. A job that must be available all day does not get this discount, which is why the usage pattern belongs in the comparison.
- What would make a second implementation a bounded cost rather than a permanent one?A written retirement date for the original and a named owner for that date. Without both, the original survives as an unfunded fallback that nobody is accountable for removing, and every change to the logic is made twice indefinitely. The difference is organisational, not technical.
- Why is validating the new implementation usually the largest part of the rewrite?Because the two programs will disagree, and each disagreement must be explained. Some differences are improvements, some are new defects, and some reveal that the original was quietly wrong for years. Separating those three requires running both over real inputs and adjudicating record by record, which is slow and cannot be skipped if anyone is to trust the result.
Running two implementations of the same logic is keeping two sets of books for the same business. Each one is perfectly sensible on its own, but every transaction now has to be entered twice, and every month somebody spends a day finding out why the two totals differ by a small amount — a day that produces no new information about the business.
saying these in an interview costs you the question
- A larger machine is just delaying the inevitable
- The rewrite's cost is the time to write the new code
- Two implementations are fine as long as both are tested
- Hardware is always more expensive than engineering time
- Keeping the original as a fallback costs nothing
- Reconciling two outputs is a one-time task at switch-over