skip to content

How do you decide which EC2 instance family and size to run a service on, and why is a fleet of many small instances not automatically cheaper or safer than fewer large ones?

level: principalimportance: should knowfreq 45%

answer

  1. measure what saturates first
  2. ratio picks the family
  3. overhead is paid per instance
  4. 'up to' bandwidth is a burst ceiling
  5. blast radius versus fixed overhead

basics

~20 s

Pick the family from the resource your workload is actually bound by, then the size from measured usage plus headroom. Small sizes carry burstable network and EBS bandwidth and pay fixed per-instance overhead repeatedly; large ones concentrate blast radius and coarsen scaling steps.

solid answer

~50 s

Start from measurement, not intuition: find which resource saturates first — CPU, memory, network, or local I/O — and let that pick the family ratio (`c` at roughly 2 GiB per vCPU, `m` at 4, `r` at 8, `i` when local NVMe throughput is the constraint). Then choose the size. The instinct that many small instances are cheaper is usually wrong for three reasons: **per-instance fixed overhead** (agents, sidecars, JVM heap floor, OS reservation) is paid on every instance; **smaller sizes get "up to" network and EBS bandwidth**, a burst ceiling over a lower sustained baseline, so throughput per vCPU is not constant down the ladder; and the per-vCPU on-demand price is roughly flat within a family, so scaling up is rarely a premium. What small instances genuinely buy is a **smaller blast radius** and a finer scaling step — losing one of twenty machines costs 5% of capacity, losing one of four costs 25%. Decide by naming the constraint that actually binds.

go deeper

for a junior

Know that the family letter should match what the workload needs most — CPU, memory or storage — and that you size from real usage rather than from a guess.

for a middle

Explain the family ratios, the fact that CloudWatch needs an agent for memory, and how per-instance overhead and burstable network bandwidth change the arithmetic between small and large sizes.

for a senior

Work through a real sizing decision end to end: identify the binding resource, choose a family and size with justified headroom, and predict the failure mode of the shape you rejected.

for a principal

Own the policy: what headroom the organisation targets and why, who revisits sizing and how often, how blast radius is traded against efficiency, and how new instance generations get adopted without a per-team project.

## Step one: find the binding constraint Right-sizing starts by identifying which resource runs out first under real traffic. Everything else follows from that answer. - CPU-bound (encoding, compilation, crypto, tight request handlers) → the `c` family, roughly 2 GiB per vCPU. - Balanced or unmeasured → `m`, roughly 4 GiB per vCPU. The honest default. - Memory-bound (caches, large heaps, in-memory joins) → `r` at roughly 8 GiB per vCPU, and the `x`/`u` families beyond that. - Local storage throughput-bound → `i` families with local NVMe. - Spiky and mostly idle → burstable `t`, accepting its baseline entitlement. The measurement itself has a trap on AWS worth naming out loud: **CloudWatch does not report guest memory utilisation by default**. CPU, network and disk come from the hypervisor, but memory requires the CloudWatch agent inside the instance. Teams that never installed it are, by construction, right-sizing on CPU alone — and then wonder why their `c` instances swap. AWS Compute Optimizer will make recommendations from the metrics you do have, and it is a reasonable starting point, but its memory advice is only as good as the agent data behind it. ## Step two: the size, and why small is not automatically cheaper **Per-instance overhead is paid every time.** Each instance carries the OS, the monitoring agent, a security agent, a log shipper, service-mesh or sidecar processes, and whatever floor your runtime needs — a JVM's non-heap footprint, a connection pool per process, a per-instance warm cache. Split 64 GiB across sixteen small machines instead of four large ones and you pay that overhead sixteen times. For memory-hungry runtimes this alone can erase the supposed saving. **Bandwidth does not scale linearly downward.** AWS quotes smaller sizes as "up to" some figure — a **burst** ceiling sustained by a network I/O credit mechanism, over a lower guaranteed baseline. The largest size in a family gets its full allocation without bursting. A fleet of small instances doing steady bulk transfer or heavy remote-storage I/O can therefore hit a sustained ceiling that its "up to" spec sheet never suggested, while a couple of large instances would not. This is one of the most common surprises in a small-instance fleet, and it is invisible in CPU graphs. **Per-vCPU price is roughly flat within a family.** Within a generation, on-demand cost scales close to linearly with size, so "scale up" is not usually a price premium — it is a different packaging of the same capacity. (How you *buy* that capacity is a separate matter.) **Some things do not shard.** A single large in-memory dataset, a licence keyed per host, a workload with a large shared cache, or anything needing more RAM than a small type offers, forces the large end regardless of preference. ## Step three: what small instances genuinely buy They are not the wrong answer — they are the right answer for different reasons than people give: - **Blast radius.** Losing one instance out of twenty removes 5% of capacity; out of four, 25%. Hardware failures, host retirements and rolling replacements all hurt less on a larger fleet of smaller machines. - **Scaling granularity.** Adding capacity in small increments tracks demand more tightly, and less capacity is wasted at each step. With four large instances, adding one is a 25% jump. - **Placement flexibility.** Smaller types are easier to place across Availability Zones and easier to find capacity for, which matters when you want a broad list of acceptable types rather than one. - **Even distribution.** Load balancing across many similar targets smooths hot spots better than across a handful. ## The synthesis The defensible position is a **middle size, several of them, across multiple Availability Zones** — enough instances that losing one is not an incident, each big enough that fixed overhead and burstable bandwidth are not dominating. Then state the exceptions explicitly: go larger when per-instance overhead is heavy, when a dataset must be resident, or when sustained network and storage throughput matter; go smaller when the workload shards cleanly and you want fine scaling steps and a small blast radius. Two further points separate a principal-level answer: **Right-sizing is a loop, not a decision.** Traffic shape changes, runtimes change, and new generations appear with better price-performance. Re-examine periodically, with an owner and a cadence, rather than treating a launch-day choice as permanent. **Headroom is a policy, not a number pulled from the air.** How much spare capacity you keep is set by how fast you can add more and how much a saturation event costs. A service that scales in two minutes can run hotter than one whose instances take ten minutes to become useful. Say which of those you are designing for, and the size follows.

  • Why is CPU-only right-sizing particularly misleading on EC2?
    CloudWatch reports CPU, network and disk from the hypervisor, but guest memory utilisation needs the CloudWatch agent installed inside the instance. Without it there is no memory data at all, so recommendations — including automated ones — optimise on CPU and can move a memory-bound service onto a compute-optimised family with half the RAM per vCPU. Install the agent before you trust any sizing exercise.
  • What does an 'up to 12.5 Gbps' network spec actually promise?
    It is a burst ceiling, not a sustained rate. Sizes quoted with "up to" use a network I/O credit mechanism: they can reach the stated figure for a while, then settle to a lower baseline. For bursty request traffic that is fine; for sustained bulk transfer or heavy remote-storage I/O it is a trap, and a larger size — or a network-enhanced `n` variant — gives a guaranteed allocation instead.
  • How does instance size interact with how quickly a fleet can absorb a traffic spike?
    Size sets the granularity of every capacity change. Large instances mean coarse steps — each addition is a big fraction of the fleet, and each new instance takes just as long to warm up while carrying more of the load once it arrives. Smaller instances track demand more smoothly and reduce the cost of overshoot, which is why fleets that scale reactively usually favour the smaller end of a workable range.
  • When is a bare-metal instance type the right choice?
    When the workload needs direct access to the physical processor: nested virtualisation, hypervisors of your own, software licensed per physical socket or core, or profiling tools that need hardware performance counters. It is not a performance upgrade for ordinary applications — current virtualised instances already run at near bare-metal speed, so choosing metal without one of those reasons just buys a large fixed unit of capacity.

saying these in an interview costs you the question

  • Assumes many small instances are always cheaper
  • Sizes purely on CPU because that is the metric available
  • Reads 'up to X Gbps' as a guaranteed sustained rate
  • Treats right-sizing as a one-time launch-day decision
  • Ignores per-instance agent and runtime overhead

context