Your company wants a multi-year AWS compute commitment, but the fleet is mid-migration to containers and part of it is moving to a different instance architecture. How do you decide what to commit to, and for how long?
answer
- Financial position against a roadmap
- Split the fleet by confidence, not by size
- Breadth beats depth during migration
- Ladder terms, buy in tranches
- Residual On-Demand is purchased optionality
basics
~20 sCommit only to the part of the baseline that survives every plausible roadmap outcome, prefer Compute Savings Plans because they follow workloads across families and into Fargate and Lambda, ladder shorter terms over the uncertain portion, and buy in tranches rather than one irreversible purchase.
solid answer
~50 sTreat the commitment as a financial position taken against a technical roadmap, and size it to the intersection of the two rather than to the bill. Concretely: find the compute baseline that exists in essentially every hour, then subtract everything the roadmap might remove, shrink or re-platform inside the term — what remains is the only spend you can honestly promise. Prefer **Compute Savings Plans** during migration, because they apply across instance families and regions and also cover Fargate and Lambda, so a workload moving from EC2 into containers or functions keeps consuming the same commitment. Ladder the terms: three years on the portion that is genuinely permanent, one year on the portion you merely expect to persist. Buy in tranches every quarter instead of one large purchase, so each is sized against observed usage rather than a forecast. And be explicit about who owns the position: the commitment is irreversible, so architecture decisions inside the term now have a finance consequence someone must be accountable for.
go deeper
Know that a commitment cannot be cancelled, so it should only cover workloads you are confident will still exist for the whole term.
Explain why a Compute Savings Plan suits a migrating fleet — it applies across families and regions and covers Fargate and Lambda — while a family-scoped plan risks being stranded.
Show the operating loop: split the fleet by confidence, commit to the durable trough, ladder terms, buy quarterly tranches, and refuse to buy while utilization is slipping.
Own the position itself. Defend a coverage target below 100% as purchased optionality, name who signs and who is accountable when the fleet changes, and make sure the commitment review sits next to the roadmap review rather than inside a finance cycle.
## The real question being asked This is not a pricing question; it is a governance question wearing pricing clothes. A multi-year commitment is an irreversible financial position — Savings Plans cannot be cancelled, exchanged or resold — taken against a roadmap that is, by construction, uncertain. The interviewer wants to know whether you can commit meaningfully without letting the commitment quietly acquire a veto over engineering decisions. ## Step one: separate the fleet into confidence tiers Stop looking at the total bill and split hourly compute usage into three tiers: - **Permanent.** Runs today, will run in three years, in recognisably the same form. Control planes, long-lived data services, the boring backbone. - **Probable.** Runs today and probably still will, but is on someone's roadmap — the services being containerised, the ones targeted for a different architecture, the region you might consolidate. - **Volatile.** Spiky, seasonal, experimental, or scheduled for deletion. Only the permanent tier deserves a three-year commitment. The probable tier gets a one-year term, where the worst case is twelve months of mismatch instead of thirty-six. The volatile tier gets nothing: it stays On-Demand, or moves to Spot if it can tolerate interruption. ## Step two: let the plan type absorb the uncertainty During a migration, plan breadth is worth more than plan depth. A Compute Savings Plan covers EC2 in any family, size, OS, tenancy and region, plus Fargate and Lambda. That is exactly the set of transitions a modernisation programme performs, so the commitment follows the workload instead of being stranded by it. An EC2 Instance Savings Plan discounts more deeply but pins the commitment to one family in one region. Its stranding risk is concrete: migrate to a different instance family and the plan matches nothing, so you pay On-Demand for the new fleet while still paying the old commitment. During a migration, buying the deeper discount is buying a slightly better rate in exchange for a materially higher chance of paying twice. Reserve the narrow plans for the permanent tier, where the family genuinely will not move. ## Step three: buy in tranches, not in one decision A single large annual purchase forces you to forecast a year of usage at one moment, usually the moment finance asks. Quarterly tranches turn one big forecast into a series of small ones, each informed by what actually happened since the last. It also produces a natural laddering effect: commitments expire on a rolling schedule rather than all at once, so no single quarter forces a decision about the entire fleet. The operating loop is simple and worth stating explicitly: 1. Measure utilization and coverage; utilization must stay essentially pegged. 2. If coverage is below target and utilization is healthy, buy a tranche against the trough of uncovered eligible usage. 3. If utilization has slipped, do not buy — find out what changed in the fleet first. 4. Repeat quarterly, and re-forecast expiries a quarter ahead so nothing lapses unnoticed. ## Step four: set a coverage target below 100 and defend it The pressure in a cost programme is always toward more coverage, because coverage looks like savings on a slide. Someone senior has to hold the line that the target is deliberately short of full coverage — a stable estate might sit in the 70–85% range, a migrating one lower — and that the residual On-Demand spend is the price of retaining the freedom to change the architecture. Framing that residual as *purchased optionality* rather than as *waste finance has not eliminated yet* is the single most useful thing a principal engineer contributes to this conversation. ## Step five: name the owner and the review An unowned commitment portfolio drifts. Decide who signs a purchase, who is accountable when utilization falls, and how a team planning a migration surfaces it before it lands. Practically that means the commitment review sits alongside the roadmap review, not inside a monthly finance meeting, so that "we are moving this fleet in Q3" reaches the person about to buy three years of it. ## What a weak answer looks like Weak answers optimise the instrument: they compare discount percentages, recommend the deepest one, and never mention that the fleet is moving — which was the entire premise. Strong answers commit less than they could, choose breadth over depth while the architecture is in motion, make the purchase incremental, and say plainly who owns the position afterwards. ## The honest tradeoff Every one of these choices costs discount. A Compute Savings Plan is cheaper than an EC2 Instance plan per dollar committed; a one-year term is worth less than three; committing to 70% of the baseline saves less than 95%. That is the point, and you should say so out loud: you are deliberately buying a smaller discount in exchange for the ability to change your mind, and the discount you gave up is the premium on that option.
- Finance pushes for 95% coverage on a three-year term. How do you argue it down?Quantify the stranding risk against the roadmap: name the workloads scheduled to move or shrink inside the term, price the commitment those changes would leave unconsumed, and compare it with the incremental discount 95% buys over 75%. Then offer the alternative — the same coverage reached in quarterly tranches as the migrations land — so the answer is "later and safely", not "no".
- How does a commitment portfolio end up constraining architecture, and how do you prevent it?It happens when a team proposes a migration and someone objects because it would strand a commitment — the billing instrument starts vetoing engineering. Prevent it by never committing to workloads that are on the roadmap to change, preferring plan types that follow the fleet, and putting the commitment review next to the roadmap review so purchases are informed by planned changes rather than surprised by them.
- What do you do about commitments already stranded by a migration that has landed?You cannot cancel them, so the choice is to consume or absorb. Look for eligible usage you could move into the plan's scope — deferring a decommission, shifting batch or non-production workloads onto the covered family and region — and if nothing fits, book it as a loss, record why it happened, and change the purchase policy that produced it rather than repeating the shape.
- Where does Spot fit into this plan?As the layer above the commitment, not as an alternative to it. Interruptible work — batch, CI, stateless scale-out behind a queue — belongs on Spot, which means it should be excluded from the baseline you commit to. Deciding that split first is what keeps you from committing to usage you intended to move to Spot anyway.
saying these in an interview costs you the question
- Sizes the commitment from total bill rather than baseline
- Buys the deepest discount while the fleet is migrating
- Makes one large annual purchase instead of tranches
- Treats residual On-Demand spend as pure waste
- Leaves nobody accountable for the commitment after purchase