skip to content

When comparing two architecture options — for example, building a custom service in-house versus buying a vendor SaaS product — how would you construct a total cost of ownership (TCO) model to make a fair comparison?

level: middleimportance: must knowfreq 80%

answer

  1. same horizon, same categories, same assumptions
  2. build's hidden cost = ongoing maintenance + on-call + opportunity cost
  3. buy's hidden cost = integration + lock-in + exit cost
  4. revisit the model periodically, not just once
  5. name your assumptions explicitly

basics

~20 s

List every cost each option will cause over the same time period — not just the price tag, but also setup, running, and people costs — then add them up and compare like for like over the same number of years.

solid answer

~40 s

A fair TCO model fixes a common time horizon (typically 3-5 years), then itemizes costs for each option across the same categories: one-time acquisition/build cost, licensing or subscription fees, infrastructure/run costs, integration and migration effort, ongoing operational labor (on-call, maintenance, upgrades), and expected costs of downtime or risk. Both options must be priced on identical assumptions — same user/traffic growth, same horizon, same discount rate if using net present value — or the comparison is misleading. The most common mistake is pricing the build option only on developer salary for the initial build and forgetting ongoing maintenance, on-call, and the opportunity cost of the team not working on something else, which systematically understates 'build' cost relative to 'buy'.

go deeper

for a junior

Should know that TCO means more than the sticker price and be able to name a couple of hidden cost categories like maintenance or support.

for a middle

Should be able to lay out a structured model with consistent categories and horizon across options, and flag the classic 'build underestimates maintenance' bias.

for a senior

Should design the model to handle uncertainty via scenarios, push back on horizon-gaming or biased assumptions, and know when NPV/discounting actually changes the recommendation.

for a principal

Should treat TCO modeling as an ongoing governance practice — revisited periodically, tied to architecture decision records, and used to justify migrations away from sunk-cost decisions across a portfolio of systems, not just a one-time build-vs-buy call.

## Why the sticker price is the wrong number Total cost of ownership modeling exists because the purchase price or the initial build estimate of an architecture option is a poor proxy for what it will actually cost the organization over its useful life. A TCO model tries to capture every dollar an option causes to be spent, directly or indirectly, from the day the decision is made until the system is retired or replaced, so that two structurally different options — say, building a bespoke internal payments service versus licensing a third-party payments platform — can be compared on the same basis rather than comparing a build estimate's sticker price against a vendor's list price. ## Building the model Mechanically, building a TCO model starts with fixing a common **time horizon**, typically three to five years for infrastructure and platform decisions, because a shorter horizon favors options with low upfront cost and a longer horizon favors options that amortize well or that lock in efficiency gains over time. Within that horizon, costs are itemized into consistent categories applied identically to every option under comparison: 1. **one-time acquisition costs** (license purchase, build effort, migration/integration work, initial training); 2. **recurring direct costs** (subscription or support fees, cloud infrastructure run costs, third-party API usage); 3. **recurring indirect costs** (the engineering time spent on ongoing maintenance, patching, upgrades, on-call incident response, and security remediation); 4. **risk-adjusted costs** (expected cost of downtime, expected cost of a security breach, cost of vendor lock-in if the vendor later raises prices or the relationship needs to be exited). Each category is estimated for every option on the same assumptions about scale, growth, and usage pattern, because comparing a build option sized for today's traffic against a buy option priced for triple that traffic is not a fair comparison. ## The costs that never appear as a line item The deeper reason this exists as a discipline, rather than just "add up the invoices," is that decision-makers systematically underweight costs that do not appear as a single line item. - **A build option's sticker price** is usually just the initial developer-hours estimate, which captures maybe 20-30% of the true lifecycle cost; the remaining cost is ongoing maintenance, the on-call burden of an in-house system with fewer eyes on it than a vendor's shared platform, the opportunity cost of engineers not working on differentiated product work, and the eventual cost of rewriting the system when its original assumptions no longer hold. - **A buy option's sticker price** is the subscription fee, but the hidden costs are integration effort, the cost of building around the vendor's limitations, data egress or per-seat pricing that scales worse than expected, and the exit cost if the vendor is later dropped. | | Build | Buy | |---|---|---| | **Sticker price** | the initial developer-hours estimate | the subscription fee | | **Hidden** | ongoing maintenance, the on-call burden, opportunity cost, the eventual cost of rewriting | integration effort, the vendor's limitations, data egress or per-seat pricing, the exit cost | A rigorous TCO model forces both of these hidden cost sets into the open on the same spreadsheet. ## The trade-off: fidelity versus speed The central trade-off in TCO modeling is **model fidelity versus decision speed**: - a model with dozens of granular line items and Monte Carlo-style uncertainty ranges is more accurate but takes analyst-weeks to build and can still be wrong if the underlying assumptions (growth rate, team velocity, vendor price increases) are wrong; - while a quick back-of-envelope model is fast but risks anchoring the decision on whichever cost happened to be easiest to estimate, typically favoring 'build' because engineers are better at estimating their own labor than at estimating five years of vendor price escalation and support costs. A well-run TCO exercise names its assumptions explicitly (e.g., "assumes 20% YoY traffic growth, no major vendor price change, 1.5 FTE ongoing maintenance load") so the model can be revisited and challenged rather than treated as a black-box number. ## Failure modes Failure modes show up repeatedly in practice: - **sunk-cost bias** where a team defends a 'build' decision made years ago even as ongoing maintenance cost has quietly exceeded what the vendor option would have cost; - **asymmetric risk pricing** where the model prices infrastructure cost carefully but assigns zero cost to the higher operational risk of an unproven in-house system versus a vendor's SLA-backed platform; - **horizon-gaming** where a proponent of one option picks a time horizon (say, one year) that happens to favor their preferred choice. A well-known real-world pattern of this is companies choosing to build in-house observability or CI/CD tooling in their early years when engineering time was cheap relative to vendor fees, then years later facing a much larger true TCO once the in-house tool required a dedicated team to maintain, at which point migrating to a vendor platform like **Datadog** or **GitHub Actions** becomes the cheaper option net of the sunk build cost — a decision that is only visible if the TCO model is revisited periodically rather than treated as a one-time exercise.

  • How do you handle uncertainty in a TCO model, such as not knowing exactly how much traffic will grow?
    Run the model under a few named scenarios — conservative, expected, and aggressive growth — rather than a single point estimate, and check whether the recommended option changes across scenarios. If the decision is only correct under the aggressive-growth scenario, that is itself an important finding to surface to stakeholders rather than hide behind a single blended number.
  • What discount rate or time-value-of-money treatment should a TCO model use, and does it usually matter?
    For most engineering decisions with a 3-5 year horizon and costs that are roughly similarly timed across options, a simple undiscounted sum is usually good enough and easier to communicate; discounting (NPV) starts to matter more when one option front-loads cost (large capex now) and the other spreads it evenly (opex over time), since money spent later is worth less in present terms.
  • Who should own building the TCO model — engineering or finance?
    It should be a joint effort: engineering owns the technical cost estimates (infrastructure, labor hours, maintenance burden) because they understand the system, while finance owns the treatment of capital costs, discount rates, and tax/accounting implications, because getting those wrong undermines the model's credibility with decision-makers.

It's like comparing buying a house outright versus renting an equivalent one: the sale price and the monthly rent are not comparable numbers by themselves — you have to add in property tax, maintenance, insurance, and the opportunity cost of the down payment to the buy side, and add in the landlord's markup and lack of equity build-up to the rent side, over the same number of years, before you can say which is actually cheaper.

saying these in an interview costs you the question

  • Compares only the vendor's list price against a rough build estimate with no maintenance line
  • Uses different time horizons or growth assumptions for the two options being compared
  • Never revisits the model after the decision is made
  • Ignores opportunity cost of engineering time entirely
  • Presents a single point-estimate number with no sensitivity or scenario range

context