skip to content

What are the most common, systemic ways that TCO and cost-benefit models for cloud architecture decisions turn out to be wrong in practice, and how would you design a modeling process to catch these failures before the decision is locked in?

level: principalimportance: should knowfreq 40%

answer

  1. operational cost >> build cost over the system's life, structurally underweighted
  2. egress + observability + add-on services are the classic missed line items
  3. lift-and-shift routinely runs 1.5-3x over estimate
  4. model = living forecast, not a one-time approval artifact
  5. state assumptions explicitly; review actuals-vs-forecast on a fixed cadence

basics

~20 s

Cost models are usually wrong in predictable ways: people forget ongoing running costs, forget data-transfer/egress fees, assume the migration will go smoothly and on time, and only check the model once instead of updating it as reality unfolds. Fixing this means building in checkpoints, naming every assumption, and revisiting the numbers after the decision, not just before.

solid answer

~50 s

The recurring, systemic failure patterns are: underestimating operational/run cost relative to the upfront build or migration cost (the '80% of TCO is the 10 years after launch' effect); ignoring cross-cutting costs like data egress, cross-region transfer, and observability/logging volume that scale with traffic in ways the core compute estimate does not capture; assuming a clean, on-time migration when 'lift and shift' projects routinely take 1.5-3x longer than planned; and treating the model as a one-time artifact used to win approval rather than a living forecast checked against actuals. A robust process addresses this by requiring every cost model to state its assumptions explicitly, building in a post-launch actuals-vs-forecast review at a fixed interval (e.g., quarterly for the first year), and assigning an owner accountable for that review, so the organization learns from its estimation errors instead of repeating them on the next project.

go deeper

for a junior

Should recognize that a cost estimate is a guess and that actual costs can turn out different from the plan.

for a middle

Should be able to name specific commonly-missed cost categories like data egress or observability cost when reviewing a model.

for a senior

Should build sensitivity analysis into models by default and push back on point-estimate-only proposals, and personally track forecast-vs-actual on their own projects.

for a principal

Should design and own the organizational process — estimation playbooks, mandatory assumption disclosure, fixed-cadence actuals reviews — that makes the whole organization's cost forecasting improve over time, not just any single model.

## What a principal is actually responsible for here TCO and cost-benefit models for architecture decisions fail in remarkably consistent, well-documented ways across the industry, and a principal-level responsibility is not just building an individual accurate model but designing the organizational process that catches these systemic failure patterns before they lock in a bad decision, and that learns from them afterward. ## Failure one: underweighting operational run cost The single most common failure is underweighting operational run cost relative to build or migration cost. The build or migration phase is visible, has a clear owner, and gets scrutinized in the approval process, while the operational cost that accrues every month for years afterward is diffuse, spread across many people's time, and rarely gets aggregated into a single number anyone is asked to defend. That operational cost is: - ongoing infrastructure spend, - on-call burden, - patching, - incident response, - incremental feature work to keep the system viable. Industry rules of thumb vary, but the pattern of the operational phase dwarfing the build phase over a multi-year system lifetime is extremely consistent, and models that anchor heavily on the build estimate systematically understate true TCO. This connects directly to the earlier build-versus-buy discussion but is a broader, systemic bias: any model that is easiest to build off the numbers that are easiest to estimate (a build quote, a license list price) will structurally under-represent the numbers that are hardest to estimate (years of operational labor, incident cost), because the difficulty of estimating something is not correlated with its actual size. ## Failure two: the cross-cutting cloud costs A second, more cloud-specific failure is ignoring cross-cutting costs that scale with usage in ways separate from the core compute or storage line item that dominates most people's mental model of "the cloud bill." These are all recurring blind spots: - **Data egress fees**, the cost of moving data out of a cloud provider's network, are notoriously easy to omit from an initial estimate because they are proportional to traffic patterns that are hard to predict pre-launch, and they can become a dominant cost for read-heavy, data-intensive, or multi-region architectures. - **Cross-region replication traffic.** - **Observability and logging pipeline costs**, which scale with request volume and log verbosity, and which teams routinely under-provision for in the estimate then over-produce in production once verbose logging is enabled for debugging and never turned back down. - **The cost of managed-service add-ons** (API gateways, service meshes, secrets managers) that get bolted onto an architecture after the original cost estimate was finalized. ## Failure three: assuming execution goes to plan A third systemic failure, specific to migration and modernization projects, is assuming execution goes to plan. "Lift and shift" cloud migrations, and monolith decomposition or replatforming efforts generally, have a well-documented pattern of running substantially over their original time and cost estimates, often cited in industry surveys and case studies at roughly one-and-a-half to three times the original projection, driven by discovering undocumented dependencies, data migration complexity, and the cost of running two environments in parallel for longer than planned. A cost model that assumes a clean six-month migration, when the organization's own historical migration track record shows twelve to eighteen months is typical, is not being appropriately calibrated against the organization's actual execution capability. ## Failure four: treating the model as a one-time artifact The fourth and most structurally important failure is treating the cost model as a one-time artifact produced to win budget approval, rather than a living forecast that gets checked against reality afterward. Once a project is approved, the incentive to revisit whether the original cost model was accurate largely disappears, especially if the project sponsor benefited from an optimistic original estimate, which means the organization never builds a track record of how wrong its own estimates tend to be and cannot calibrate future estimates against that history. This is the same discipline that separates mature engineering-estimation cultures from immature ones in any domain: an organization that never checks estimates against actuals cannot get better at estimating. ## Designing the process that catches them Designing a process to catch these failures before and after the decision locks in requires a few concrete structural elements. 1. **Before approval**, every cost model should be required to state its assumptions explicitly and separately from its conclusions (traffic growth rate, migration duration, discount/reservation assumptions, whether egress and observability costs were modeled), so a reviewer can challenge the assumptions directly rather than just the bottom-line number. 2. **Run it under scenarios.** The model should be run under at least a conservative and an aggressive scenario, not a single point estimate, specifically to surface how sensitive the recommendation is to the riskiest assumptions (usually migration duration and growth rate). 3. **After approval**, the process should mandate a fixed-cadence actuals-versus-forecast review, for example quarterly through the first year post-launch, with a named owner accountable for producing it, and that review's findings should feed back into an organizational estimation playbook ("our lift-and-shift projects have historically run 1.7x over estimate; apply that multiplier going forward") so estimation accuracy compounds over time rather than resetting to zero on every new project. A well-known real-world instance of this discipline operating at scale is how mature FinOps and cloud-center-of-excellence practices at large cloud-native enterprises maintain running post-mortems on major migration and infrastructure investments specifically to build an internal calibration dataset, treating cost estimation as a skill the organization improves deliberately rather than a one-off exercise repeated from scratch, in the same spirit as how mature software organizations track estimation accuracy on sprint velocity to calibrate future planning.

  • How would you build an internal 'estimation accuracy' track record without it becoming a blame exercise that discourages honest forecasting?
    Frame the actuals-vs-forecast review around calibrating the process and its inputs, not judging the individual who made the estimate, and publish the resulting adjustment factors (like a migration-duration multiplier) as reusable organizational knowledge; making clear that the review exists to improve future estimates, not to score past ones, is what keeps people forecasting honestly instead of sandbagging or padding defensively.
  • Data egress cost is notoriously hard to estimate pre-launch. What practical technique would you use to get a defensible estimate anyway?
    Instrument a proof-of-concept or a comparable existing workload to measure actual outbound data volume per unit of traffic, then apply the provider's egress pricing to that measured rate rather than guessing, and explicitly flag egress as a line item with wider-than-usual uncertainty bounds in the model so reviewers know it deserves extra scrutiny as usage scales.
  • If a migration's actuals come in significantly over the original forecast, what should happen to the projects that were approved based on the original, now-known-to-be-wrong ROI model?
    The project should be re-evaluated with the actual cost-to-date and a revised remaining-cost estimate to check whether the original business case still holds going forward, since sunk cost should not drive the decision to continue, and the lesson about the estimation error should be captured for the organizational estimation playbook regardless of whether the project continues or is cut.

It's like estimating the cost of adopting a puppy by pricing only the adoption fee: the real, dominant cost is years of food, vet bills, and time, which is exactly the part people underestimate because it's diffuse and spread over years instead of being one visible number on day one — and if you never go back and check what you actually spent in year one against what you guessed, you'll make the same underestimate with the next puppy.

saying these in an interview costs you the question

  • Presents a single point-estimate TCO number with no stated assumptions or sensitivity range
  • Cost model has no line item for data egress, observability, or add-on managed services
  • Assumes a migration will complete on the original schedule with no reference to the organization's own migration track record
  • No plan to check the model against actuals after the project ships
  • Treats an estimation miss as a one-off surprise rather than a pattern to learn from

context