For an EC2 workload, how do you decide whether data belongs on an instance store volume rather than an EBS volume, and what must be true of the application before you choose instance store?
answer
- can this be rebuilt without it?
- latency versus every durability guarantee
- local I/O bypasses the EBS bandwidth ceiling
- replication factor, not the disk, is the durability
- most real designs use both
basics
~20 sChoose EC2 instance store only when the data is reproducible — scratch, spill, cache, or a replica of a shard held elsewhere — and the workload is limited by storage latency or throughput. Everything that is a source of truth belongs on EBS or a managed service.
solid answer
~50 sThe decision is one question asked twice. First: if this host disappears without warning, can the data be rebuilt — from S3, from a database, from peer replicas, or by recomputation? If the answer is no, it is EBS or a managed store, and the conversation ends. Second: does the workload actually need what instance store gives you? Local NVMe sits inside the host, so its latency is not a network round trip and its I/O does not consume the instance's EBS bandwidth allocation — that matters for spill-heavy analytics, high-rate local caches, media scratch, and replicated datastores, and is irrelevant for a service that is CPU-bound or does a handful of IOPS. It is also bundled into the instance price rather than billed per GB and per provisioned IOPS. If both answers are yes, use it; if only the second is, you are trading durability for speed you cannot pay for.
go deeper
Be able to say instance store suits temporary data — scratch, caches, working files — and that anything you cannot afford to lose goes on EBS or a managed service.
Explain the mechanics behind the choice: local disks avoid the network hop and do not consume the instance's EBS bandwidth allocation, but come with no persistence, snapshots, or independent sizing.
Show the decision procedure and the recovery arithmetic: what a lost node costs, how long a rebuild from replicas takes, and whether a cold cache stampedes the origin. Mention that most designs split state across both volume types.
Own the fleet-wide standard for which data classes may live on ephemeral storage, and the cost comparison that keeps teams from buying storage-optimised instance types for workloads that never touch the disks.
## The tradeoff in one line Instance store gives you the fastest storage EC2 offers and takes away every durability guarantee. EBS gives you a volume that outlives the instance and can be snapshotted, moved, and re-attached, at the cost of a network hop and a separate bill. Everything else is detail. ## What instance store actually buys **Latency and IOPS without a network hop.** The disks are physically in the host. There is no fabric between the instance and the device, which is why local NVMe latency is dramatically lower than network-attached block storage under the same queue depth. **Bandwidth that is not your EBS bandwidth.** On Nitro instances, EBS traffic runs over a dedicated allocation that scales with instance size and is a real ceiling for I/O-heavy work. Local NVMe I/O does not consume it. A workload that spills tens of gigabytes to disk can saturate its EBS allocation and starve everything else on the instance; moving spill to instance store removes that contention entirely. **No separate storage bill.** Local capacity is included in the instance's hourly price. There is no per-GB-month charge and no provisioned-IOPS or provisioned-throughput charge, and stopped-instance storage cost does not apply because you cannot usefully stop the instance anyway. ## What it costs Everything the sibling technology provides: persistence across stop/start, survival of host failure, snapshots, restore, detach and re-attach, resize, and capacity chosen independently of the compute shape. Local capacity is fixed by the instance type, so you size compute and storage together whether or not that ratio suits you. ## The gate: is the data reproducible? Run the thought experiment honestly. The host is retired at 03:00 with no notice. What is lost, and what does recovery cost? - **Nothing is lost** — the data was scratch, spill, temp, or intermediate state for an in-flight job that will be retried. Instance store is the right answer. - **A cache is lost** — the service warms back up from its real source of truth. Fine, provided the cold-start stampede against that source is survivable. This is a capacity question, not a durability question. - **One replica of a shard is lost** — a datastore with a replication factor above one rebuilds that node from its peers. Fine, provided you have modelled the rebuild: how long it takes, how much network it consumes, and whether you can lose a second node during it. - **The only copy is lost** — stop. This is a system of record. It belongs on EBS, S3, or a managed database, whatever the performance argument. ## Workloads that genuinely fit - Analytics and ETL engines that spill shuffle and sort data to local disk. - Build agents and CI workers with large, disposable working trees and dependency caches. - Media transcode and rendering scratch, where the source is in S3 and the output is written back. - Local caches in front of a database or object store. - Distributed datastores and search clusters that own their own replication and are designed for node replacement. ## Workloads that do not - A single-writer database primary holding the system of record. - Anything one instance writes and another must read — that is a shared-filesystem problem, not a local-disk one. - A boot volume you want to snapshot and roll back; boot volumes on modern instances are EBS. - Long-lived state on a fleet you routinely stop overnight to save cost — the first stop erases it. ## A hybrid is usually the real answer Most production designs use both on the same instance: EBS for the root volume and any state that must survive, instance store for the hot path. A search node keeps its data directory on instance store and relies on cluster replication; a build host keeps `/` on EBS and its workspace on local NVMe. Splitting by durability requirement rather than picking one volume type for everything is the answer that reads as experience. ## The cost trap in reverse Instance store looks free because there is no line item. It is not: you pay for it in the instance type. A `d`-suffixed or storage-optimised instance costs more per hour than its plain equivalent, so choosing local NVMe for a workload that does a handful of IOPS means paying for disks you never touch. Compare the fully-loaded instance-hour plus zero storage against the plain instance-hour plus the EBS you would have provisioned, and be sure the latency actually shows up in a metric you care about.
- Your service is I/O-bound and you have already moved it to instance store. How do you show the change actually helped?Compare storage-side latency and queue depth before and after, not just throughput. On EBS you would also watch whether the instance was hitting its EBS bandwidth or IOPS allocation — a workload pinned against that ceiling is the clearest case for local NVMe. If latency percentiles and the EBS ceiling both look untouched, the workload was never storage-bound and you bought durability loss for nothing.
- A team wants to run their primary PostgreSQL on instance store for the IOPS. What is your answer?No for a single primary — a host retirement destroys the system of record with no restore path. If the IOPS requirement is real, the options are a higher-performance EBS volume type, or an architecture where the local disk is not the only copy: streaming replication to a node on durable storage plus continuous archiving, so the local disk is a performance layer and not the source of truth.
- How does the fixed compute-to-storage ratio of these instance types affect capacity planning?You cannot scale local storage independently, so the instance type is chosen by whichever dimension binds first. If you need more local terabytes you buy more vCPU and memory with them, and vice versa. Where the ratio is badly wrong for the workload, EBS is often cheaper overall precisely because capacity and compute are decoupled.
saying these in an interview costs you the question
- Choosing instance store purely for speed on critical data
- Assuming instance store is free because there is no line item
- Treating a replicated cluster as immune to node loss
- Ignoring that the local capacity is fixed by instance type
- Putting a single-writer system of record on ephemeral disk