skip to content

Beyond the doubled infrastructure bill, which recurring costs appear when a workload is kept live on two cloud platforms?

level: seniorimportance: should knowfreq 47%

answer

  1. the infrastructure line is the small half
  2. two of everything not the application
  3. capacity bought twice, mostly idle
  4. split spend, smaller commitment discount
  5. either widen the rota or train everyone twice

basics

~20 s

The lasting costs are people and process: two control surfaces, two identity models, two quota regimes, two audits and two incident procedures. Add capacity bought twice so either side can carry the load alone, a split commitment discount, and continuous replication traffic charged on the way out.

solid answer

~40 s

The infrastructure line is the least interesting part of the tax. Operating two platforms means **two of nearly everything that is not the application**: management APIs, identity models, quota and throttling behaviour, audit trails, maintenance calendars and incident procedures. On-call is the sharpest edge — you either widen the rota or require every engineer to be genuinely competent on two platforms, and competence means having handled incidents there, not having read about it. Capacity is bought roughly twice, since each side must carry the whole load alone. Your spend is split, so each supplier sees a smaller volume and you qualify for a smaller commitment discount than one platform would have given. And keeping the two copies in step generates continuous traffic that is billed leaving the platform. None of these decrease over time.

code

pseudocode · 29 lines
pseudocode
# illustrative only - every figure below is invented and qualitative

baseline = cost_of_workload_on_one_platform_per_year

# capacity: each side must be able to carry the whole load alone
compute = baseline.compute * 2

# commitments: spend is split, so each supplier grants a weaker discount tier
discount_loss = baseline.commitment_discount * split_penalty

# tooling, controls, audit evidence: built and evidenced on both platforms
platform_tooling = baseline.tooling * 2

# people: either a wider rota, or every engineer competent on two platforms
if rota_strategy == "widen":
    people = baseline.oncall + cost_of_additional_rota_members
else:
    people = baseline.oncall + cost_of_dual_training_for_every_engineer

# keeping two live copies in step is continuous traffic, billed on the way out
sync_traffic = gigabytes_replicated_per_month * 12 * per_gigabyte_rate

total = compute + discount_loss + platform_tooling + people + sync_traffic
tax   = total - baseline.total

if tax > value_of_surviving_the_loss_of_a_whole_platform:
    decision = "posture does not earn its tax for this workload"
else:
    decision = "posture is justified for this workload only"

go deeper

for a junior

Recall that the second infrastructure bill is not the main cost. Operating two platforms means two management interfaces, two identity models, two audit trails and two incident procedures, every year.

for a middle

Explain why the unit cost rises as well as the unit count: splitting spend across suppliers earns a weaker commitment discount on each side, and each side must be sized to carry the whole load alone.

for a senior

Show the operating reality — which limits differ, which audit evidence doubles, and how the rota is staffed so that whoever is paged has actually handled an incident on that platform rather than read about it.

for a principal

Treat the competence and controls cost as a permanent line in the budget and confine the posture to the tier that justifies it. Decide explicitly whether the second side is full-capacity or degraded-mode, and make that decision visible.

## The bill everyone models, and the bill that actually hurts When a board asks what a second platform costs, the answer offered is usually the infrastructure line. That is the easiest number and the least important one. The live-on-both posture charges four distinct recurring taxes, and only the first is what people mean by "the cloud bill". ## The four taxes 1. **Capacity bought twice.** The posture only delivers what it promises if either side can carry the whole load alone. That means roughly two full-size estates, both running, both mostly idle. Sizing the second side for half the load is a common compromise and it silently converts the posture into degraded-mode survival — legitimate, but it must be stated, because the failure behaviour is no longer what was sold. 2. **A split commitment discount.** Term commitments are discounted because you promised a supplier a volume. Halving your spend across two suppliers means each sees a smaller promise, so the discount tier you qualify for on each side is worse than the one your undivided spend would have earned. The effective cost per unit rises on both platforms at once. 3. **Doubled operating surface.** This is the tax that never shrinks: - two management APIs, each with its own throttling behaviour and its own asynchronous provisioning quirks; - two identity models, and a mapping between them that somebody owns; - two quota and limit regimes, with different processes for raising a soft quota; - two audit trails to collect, retain and evidence to an auditor; - two maintenance and deprecation calendars that do not align; - two sets of controls, each needing its own evidence at review time. 4. **Traffic to keep the copies in step.** Two live copies of a stateful system exchange data continuously, and traffic leaving a platform is billed. The rate schedule itself is a separate subject, but the shape matters here: this is a standing cost that scales with write volume, not a one-off. ## The people tax deserves its own heading The hardest recurring cost is competence. During an incident the engineer on call must know how *this* platform signals a degraded management API, where *this* platform's audit trail lives, and which of *this* platform's limits is about to be hit. That knowledge decays when it is not used. There are two honest ways to hold it, and both cost: - **Widen the rota** so specialists for each platform are reachable — more people carrying a pager, for the same volume of service. - **Make everyone dual-competent** — real training time, real practice on both platforms, and slower onboarding for every new hire, permanently. The unstated third option, *hope the person on call happens to know*, is how a platform-specific incident turns into an hour of hunting for the right console. ## Recomputing the tax honestly A quick model, with every figure invented and qualitative, shows why the infrastructure line misleads. Take a workload costing one unit of infrastructure a year on a single platform. Under the live-on-both posture: compute and storage roughly double to two units; the lost discount tier adds a fraction of a unit on top of that; tooling, audit evidence and controls work roughly doubles; and on-call either grows the rota or grows training. The infrastructure part of the increase is the part you can see in a bill, and it is frequently the smaller half of the total once people and process are counted. ## What follows from this - **Scope the posture, do not apply it estate-wide.** Pay the tax on the tier whose unavailability is genuinely intolerable, and let everything else sit on one platform. - **State the degraded-mode decision explicitly.** If the second side is sized for half the load, say so, and say what is shed when it takes over. - **Budget the people cost as a standing line**, not as a project cost. It appears every year, for every new hire, and it is the first cost that quietly gets skipped — which is how an estate ends up holding a second platform that nobody on call can actually operate.

  • Why does splitting spend across two suppliers make each unit of capacity more expensive?
    Commitment discounts are bought with a promised volume. Two half-sized promises land in worse discount tiers than one whole-sized promise, so the effective rate rises on both platforms simultaneously. The posture therefore raises unit cost in addition to doubling the units, which is why a naive model that simply multiplies today's bill by two understates it.
  • A team sizes the second platform for half the peak load to save money. What has changed about the posture?
    It is no longer full survival; it is degraded-mode survival. That can be a perfectly good decision, but it must be written down along with what gets shed — which customers, which features, which batch work — and the shedding has to be tested. The failure mode that is never acceptable is discovering the sizing choice during the outage.
  • Which of these recurring costs shrinks as the team gets better at the posture?
    Some of the tooling and process cost does, because automation and shared abstractions absorb repeated work. Capacity, the weakened discount tier and the replication traffic do not — they are structural. Neither does the competence cost, because it resets with every new hire and decays whenever a platform goes a long stretch without an incident.

It is a second fully staffed kitchen kept hot so the restaurant can serve if one burns down. Two rents, two crews, two inspections — and the menu shrinks to dishes both kitchens can cook.

saying these in an interview costs you the question

  • Modelling the tax as simply doubling today's infrastructure bill
  • Forgetting that split spend earns a weaker commitment discount
  • Assuming one on-call engineer can operate both platforms untrained
  • Ignoring the continuous replication traffic billed leaving each platform
  • Sizing the second side for half the load without saying so
  • Treating the people and audit cost as a one-off project expense