skip to content

An EC2 instance's Amazon EBS volume is provisioned for 16,000 IOPS, but a benchmark tops out near 6,000 with high latency. What AWS-side causes would you investigate, and how does that shape your volume-type choice?

level: seniorimportance: nice to knowfreq 38%

answer

  1. the volume is not the only limit
  2. the instance has a ceiling too
  3. large I/O is counted in chunks
  4. depth one never reaches the number
  5. fresh from a snapshot means cold blocks

basics

~20 s

Suspect the instance before the volume: every EC2 instance type has its own EBS bandwidth and IOPS ceiling, and smaller sizes only burst to it. Then check I/O size against the throughput cap, queue depth, and whether the volume was restored from a snapshot and is still lazily loading blocks.

solid answer

~50 s

Provisioned IOPS is a ceiling on the volume, not a guarantee on the path. The first thing I check is the **instance**: each EC2 instance type publishes its own maximum EBS bandwidth and IOPS, some sizes only burst to it rather than sustaining it, and that limit is shared across every attached volume — so a big volume on a small instance is throttled at the instance. Second, **I/O size**: EBS counts SSD I/O in 256 KiB units, so large requests consume multiple IOPS and you may be hitting the throughput cap rather than the IOPS one; on gp3 throughput is also bounded relative to provisioned IOPS. Third, **queue depth** — a single-threaded synchronous workload simply cannot keep enough I/O outstanding to reach 16,000. Fourth, if the volume came from a snapshot, first-touch reads are still being pulled from snapshot storage. Only once those are ruled out does moving to io2 Block Express make sense.

code

bash · 4 lines
bash
# generate enough outstanding I/O to actually reach the provisioned rate
fio --filename=/dev/nvme1n1 --rw=randread --bs=16k \
    --iodepth=64 --numjobs=4 --ioengine=libaio --direct=1 \
    --runtime=60 --time_based --group_reporting --name=ebs

go deeper

for a junior

Know that provisioned IOPS is an upper limit on the volume, not a promise about what any given application will achieve end to end.

for a middle

Explain the other limits: the instance type's own EBS bandwidth, I/O counted in 256 KiB units on SSD volumes so large requests cost several IOPS, and gp3 throughput bounded by provisioned IOPS.

for a senior

Show the diagnostic order — aggregate the instance's volumes against its documented ceiling, check queue depth and I/O size, rule out snapshot lazy loading — and only then change volume type.

for a principal

Own the guidance that stops the organisation buying performance it cannot use: measure the binding constraint first, standardise instance and volume pairings, and decide when a workload should be spread across instances instead of scaled up.

## The mental model: three ceilings, not one Between an application and its data sit three independent limits, and the smallest wins: 1. **The volume's provisioned performance** — what you bought. 2. **The instance's EBS bandwidth and IOPS ceiling** — a property of the instance type and size, shared across all its EBS volumes. 3. **The workload's own concurrency** — how much I/O it manages to have in flight. Candidates who only know the first spend a lot of money raising a number that was never the constraint. ## Ceiling one: the instance Every current instance type publishes a maximum EBS bandwidth (MiB/s) and a maximum EBS IOPS figure. Two things about it catch people out. It is **aggregate**: attach four volumes and they share it, so a striped set does not multiply past the instance limit. And on many smaller sizes it is a **burst** figure — the instance can hit the headline number for a while and then settles to a lower baseline, which looks exactly like a volume problem if you are only watching volume metrics. The diagnosis is to compare observed aggregate throughput and IOPS against the instance type's documented EBS limits, and to test the same volume from a larger instance. If the number moves with the instance size, you have your answer, and the fix is a bigger or more EBS-optimised instance rather than more provisioned IOPS. ## Ceiling two: I/O size and the throughput cap EBS does not count "an operation" naively. For SSD-backed volumes an I/O is measured in units of up to **256 KiB** — a 1 MiB request consumes four IOPS. For HDD volumes the unit is **1 MiB**, and sequential I/O is merged, which is exactly why `st1` is good at scans and dreadful at random access. So a workload doing large sequential reads will exhaust the volume's **throughput** limit long before its IOPS limit, and the CloudWatch picture looks like "IOPS well under the provisioned number, yet everything is slow". On `gp3`, throughput is additionally bounded relative to provisioned IOPS (roughly 0.25 MiB/s per IOPS, with a 1,000 MiB/s per-volume ceiling as of 2025), so raising throughput sometimes means raising IOPS you do not otherwise need. The inverse trap is a workload of tiny random reads: it will hit the IOPS ceiling with the throughput graph nearly flat. ## Ceiling three: queue depth A volume cannot serve 16,000 IOPS to a client that only ever has one request outstanding. If average latency is 0.5 ms, a single-threaded synchronous reader achieves about 2,000 IOPS and no more — not because anything is throttled, but because it never asks for more. Reaching high provisioned rates requires deep queues: many concurrent threads, asynchronous I/O, or a database configured with enough parallel I/O workers. A benchmark should say so explicitly. Running with a queue depth of 1 and concluding the volume is slow is the single most common false alarm: ```bash fio --filename=/dev/nvme1n1 --rw=randread --bs=16k \ --iodepth=64 --numjobs=4 --ioengine=libaio --direct=1 \ --runtime=60 --time_based --name=ebs ``` Watch the volume's queue length in CloudWatch alongside latency: a deep queue with high latency means you are at a real limit; a shallow queue with high latency points at the client or the path. ## The snapshot restore case If the volume was created from a snapshot, untouched blocks are fetched from snapshot storage on first read. A benchmark over a freshly restored volume measures lazy loading, not steady-state performance. Warm the device by reading it through once, or use Fast Snapshot Restore for the snapshot in that AZ, then re-measure. ## What this implies for volume-type choice Once you know which ceiling binds, the choice follows: - **Random IOPS is the constraint and 16,000 is genuinely not enough** — `io2`, and `io2 Block Express` for the top end, which also brings more consistent sub-millisecond latency and a higher stated durability. Provisioned IOPS types are also the only ones supporting Multi-Attach. - **Sustained sequential throughput on large data** — `st1` (throughput-optimised HDD) is far cheaper per GiB and built for streaming; `sc1` for colder, rarer scans. Neither is a boot volume and neither tolerates random access. - **Everything else** — `gp3`, sized for capacity and dialled for the IOPS and throughput you measured. - **The instance is the wall** — change the instance, or spread the workload across more instances; no volume type fixes it. The answer an interviewer is listening for is that you measure before you buy, and that you know performance is a property of the whole path rather than a number on the volume.

  • Would striping four EBS volumes together give four times the IOPS?
    Only up to the instance's own EBS ceiling, which is shared across all attached volumes. Striping helps when the per-volume limit binds and the instance still has headroom; it does nothing once the instance is the constraint. It also multiplies failure and snapshot-coordination complexity, so it is a last resort, not a default.
  • How do you tell an instance-level throttle from a volume-level one in CloudWatch?
    Sum throughput and IOPS across every volume on the instance and compare that against the instance type's documented EBS limits — a volume sitting below its own provisioned figure while the aggregate sits at the instance figure is an instance throttle. Repeating the test on a larger instance size confirms it quickly.
  • Your monitoring shows IOPS far below the provisioned number but throughput pinned. What does that indicate?
    Large I/O requests. SSD volumes count I/O in units of up to 256 KiB, so big sequential requests burn throughput while the operation count stays low. You are at the throughput ceiling, not the IOPS ceiling — the fix is more provisioned throughput, or a throughput-optimised type such as st1 if the pattern is genuinely sequential.

saying these in an interview costs you the question

  • Assumes provisioned IOPS is guaranteed regardless of instance type
  • Benchmarks at queue depth one and blames the volume
  • Ignores that instance EBS bandwidth is shared across all volumes
  • Measures a volume freshly restored from a snapshot as steady state
  • Reaches for io2 before finding which ceiling actually binds

context