skip to content

You lead a platform team where every product team picks its own AWS compute model. How do you decide on a default compute model for the organisation, and when do you let a team deviate?

level: principalimportance: nice to knowfreq 36%

answer

  1. count the roads, not the services
  2. fixed cost per supported model
  3. default serves the median workload
  4. exceptions on attributes, not preference
  5. make the road the easiest path

basics

~20 s

Pick the default that fits the median workload and that your team can actually pave — deployment, observability, identity, on-call. Allow deviation only when a workload has an attribute the default cannot serve, and make that a documented, reviewed decision rather than a preference.

solid answer

~50 s

I would optimise for the total cost of the *paved road*, not the compute line. Every supported compute model needs its own deployment path, log and metric conventions, identity pattern, runbooks and on-call knowledge, and that fixed cost is paid per model, per year, whether or not anyone uses it. So I would inventory what teams actually run, choose the default that covers the median workload well, and invest deeply in that one road so it is genuinely the easiest thing to do. Deviation is allowed on workload attributes the default cannot serve — a hard duration ceiling, a GPU, a protocol, a latency commitment, licensing — never on preference or novelty. I would make the exception path real but explicit: a short written case, a named owner, and a review date, so exceptions stay countable. And I would measure the road's adoption rate, because if teams keep leaving it, the road is the problem, not the teams.

go deeper

for a junior

Know that organisations usually standardise on one compute model, and that the reason is consistency in deployment, monitoring and on-call rather than the price of compute itself.

for a middle

Explain what a paved road contains — pipeline, observability, identity, runbooks — and why each additional supported model duplicates all of it.

for a senior

Show how you would run the exception path in practice: attribute-based criteria, a written owner and review date, and evidence gathered from a real workload inventory.

for a principal

Own the tradeoff explicitly, name the metrics that tell you whether the standard is working, and be able to say what standardisation costs the teams it does not suit.

## Reframe the question Asked as "which compute model is best", this has no answer. Asked as "how many compute models can we afford to be good at", it does. The cost of a compute model to an organisation is not its hourly price; it is the paved road that has to exist around it before a product team can ship safely: - a build and deploy path, with rollback - log, metric and trace conventions that make dashboards comparable across services - an identity pattern — how a workload gets credentials without static keys - network placement and egress conventions - runbooks, alarms and enough shared on-call familiarity that someone other than the author can debug it at 3 a.m. - security review, guardrails and cost attribution That list is paid once per model and then maintained forever. Two models is not twice one model's compute bill; it is roughly twice that fixed organisational cost, and it dilutes the operational familiarity that makes incidents short. ## Choosing the default Start from evidence: inventory what teams run today, what shapes those workloads have, and where incidents and delivery friction actually come from. Then choose the model that serves the *median* workload well, not the most interesting one. In most product organisations that is containers — they fit the widest range of workloads, impose the fewest constraints on how code is written, and keep the local development story close to production. Organisations dominated by event-driven glue, low-duty-cycle internal tools and integrations often land on functions instead, and that is equally defensible. Two weightings deserve to be explicit. **Team shape**: a default that needs a platform team you do not have is not a default, it is a wish. **Reversibility**: a default that makes it cheap to move a service later is worth more than one that is marginally better today, because you will get this wrong for some workloads. ## Designing the exception path The failure mode at both extremes is well known. Too rigid, and teams route around the platform — you get shadow infrastructure with none of the guardrails. Too permissive, and you have five models, five half-built roads and no depth in any. So make deviation legitimate but *costly to do casually*: 1. **Criteria, not vibes.** An exception is granted on a workload attribute the default genuinely cannot serve — a duration ceiling, a GPU or specialised hardware, a protocol or port the default does not expose, a licensing requirement, a latency commitment the default cannot meet, or a regulated isolation requirement. "The team prefers it" and "it's what we know" are not attributes of the workload. 2. **Written and owned.** A short document: the attribute, the alternative considered, who owns the deviation, and a review date. 3. **Time-boxed.** Revisit at the review date. Platforms change; a constraint that was real two years ago may not be. 4. **Counted.** Track how many exceptions exist. A steadily rising count is a signal that the default is wrong or the road is bad. ## Make the default the path of least resistance Standards enforced by policy alone decay. The road wins by being easier: a template that gives a new service a pipeline, logging, dashboards, alarms, identity and cost tags on day one. If choosing the default saves a team a week and choosing something else costs them that week, you rarely need to argue. This is also why a migration mandate is usually the wrong first move. Existing services that work are not the problem; the flow of *new* services is. Set the default for new work, migrate opportunistically when a service is being changed substantially anyway, and let the population converge. ## What to measure Name the metrics, because that is what separates a principal answer from an opinion: share of new services launched on the default; time from empty repository to production; number and age of open exceptions; incident time-to-mitigate split by model; and the fraction of compute spend on capacity that is idle. Those tell you whether the standard is real, whether the road is good, and whether the default is still the right one. ## The tradeoff to state out loud Standardisation trades local optimality for organisational leverage. Some teams will run on a model that is not the best fit for their workload, and they are right that it is suboptimal for them. The counter-argument is that the organisation gains faster onboarding, comparable observability, transferable on-call skill and fewer ways to get security wrong — and that those compound while the local inefficiency does not. A leader who cannot articulate what standardisation *costs* the dissenting team has not actually made the tradeoff, only imposed it.

  • A team argues the default costs them 30% more in compute for their workload. How do you respond?
    Take the number seriously and compare it against what the exception costs the organisation — tooling, review, on-call familiarity, and the precedent. If 30% of a small bill is a few hundred dollars a month, the default holds; if it is a material line item on a large service, that is exactly the evidence an exception process exists to act on.
  • How do you avoid the default becoming a ceiling as the company changes?
    Give it a review cadence and let the exception log drive it. Exceptions clustering around one attribute are telling you the default no longer covers the median workload. Treat a rising exception count as a platform defect rather than as teams misbehaving.
  • Would you migrate existing services onto the new default?
    Not as a campaign. Set the default for new services, migrate opportunistically when a service is already being reworked, and leave stable systems alone. A migration mandate spends real engineering capacity on services that were not causing problems, and it is the fastest way to lose credibility for the platform.

saying these in an interview costs you the question

  • Mandates one model with no exception path at all
  • Chooses the default from personal preference rather than the workload inventory
  • Ignores that each supported model needs its own tooling and on-call knowledge
  • Orders a mass migration of working services before fixing the paved road
  • Cannot say what standardisation costs the teams it does not suit

context