skip to content

A production EC2 instance's Amazon EBS data volume is nearly full. How do you grow it without downtime, and what are the limits of an EBS volume modification?

level: middleimportance: should knowfreq 52%

answer

  1. the API resizes the device only
  2. the OS still has to be told
  3. growth is one-directional
  4. a cooldown gates the next change
  5. snapshot before you touch partitions

basics

~20 s

Call ModifyVolume with the new size while the volume stays attached, wait for it to leave the optimizing state, then extend the partition and filesystem inside the operating system. EBS volumes can only grow, never shrink, and a cooldown applies before the next modification.

solid answer

~50 s

Growing an EBS volume is an online operation. I call `ModifyVolume` on the attached volume with a larger `--size` (and optionally a new type, IOPS or throughput in the same call), then poll `DescribeVolumesModifications` — the volume moves through `modifying` and `optimizing` and stays readable and writable the whole time, though performance during optimisation sits between the old and new settings. AWS only changes the block device; the guest OS still sees the old partition and filesystem size, so the second half of the job is extending the partition and then growing the filesystem in place. The hard limits worth naming: a volume can only ever be enlarged, and there is a cooldown before the same volume can be modified again — six hours at the time of writing — so bumping a volume by 10 GiB repeatedly will lock you out. Size generously.

code

bash · 7 lines
bash
aws ec2 create-snapshot --volume-id vol-0abc --description "pre-resize"

aws ec2 modify-volume --volume-id vol-0abc --size 500

# poll until the modification leaves the modifying state
aws ec2 describe-volumes-modifications --volume-ids vol-0abc \
  --query 'VolumesModifications[0].[ModificationState,Progress,TargetSize]'

go deeper

for a junior

Know that an EBS volume can be enlarged while the instance keeps running, and that the operating system needs its own step afterwards before the extra space is usable.

for a middle

Explain the modifying and optimizing states, that the same call can change type, IOPS and throughput, and that volumes only grow — with a cooldown before the next modification.

for a senior

Show the safe runbook: snapshot first, size for months not days because of the cooldown, extend partition and filesystem, then verify free space and volume latency in monitoring rather than declaring victory at the API call.

for a principal

Own the capacity policy — default volume sizing, alarm thresholds that leave room to act, and a clear view on which datasets should not live on a single growable disk at all.

## The operation has two halves The single most common mistake is doing only the AWS half and wondering why `df` is unchanged. Growing storage under a running instance is always two steps: **AWS resizes the block device**, then **the guest OS grows the structures on top of it**. ### Half one: Elastic Volumes EBS Elastic Volumes lets you change a volume's size, type, IOPS and throughput **while it is attached and in use**, on current-generation (Nitro) instances, with no detach, stop, or reboot. One API call carries all the dials you want to change: ```bash aws ec2 modify-volume --volume-id vol-0abc \ --size 500 --volume-type gp3 --iops 6000 --throughput 250 ``` The volume then reports a modification state you can poll with `DescribeVolumesModifications`: - `modifying` — the change has been accepted and is being applied. - `optimizing` — the new configuration is live but the volume is still being rebalanced in the background; it is fully usable, at performance somewhere between the old and new settings. - `completed` — done. This works on the **root volume** too, which surprises people; you do not need to stop the instance to grow it on modern instance types. ### Half two: partition and filesystem Until you act inside the instance, the operating system still sees the old geometry. If the volume is partitioned, the partition table must be updated to claim the new space, and then the filesystem grown into the enlarged partition. Both are standard Linux operations that work on a mounted, live filesystem for the common filesystem types; the specifics belong to the OS, not to EBS. What matters for the AWS answer is that **the step exists and is yours to run** — nothing about `ModifyVolume` reaches inside the guest. ## The limits that actually bite **You can only grow.** There is no shrink operation, and no way to restore a snapshot into a smaller volume. If you genuinely need a smaller disk, you create a new, smaller volume, copy the data across at the filesystem level, and swap the attachment — which is real downtime and real risk. This is the reason to think about growth before you provision, not after. **There is a modification cooldown.** After a successful modification, the same volume cannot be modified again for a waiting period — six hours at the time of writing. Teams that automate "add 10 GiB when 90% full" discover this the first time the volume fills twice in an afternoon and the second call is rejected. Grow in meaningful increments, and alarm early enough that you are not resizing under pressure. **Throughput during optimisation is not the new number yet.** Do not schedule a modification as the fix for a load spike already in progress and expect instant relief. **The instance has its own ceiling.** Enlarging a volume or raising its provisioned IOPS does nothing if the instance type's EBS bandwidth is already the bottleneck. Check both sides. ## Doing it safely A sensible production sequence: 1. Take a snapshot first. It is incremental and cheap, and it is your undo button if the partition step goes wrong. 2. `ModifyVolume` with the new size — and size for the next twelve months, not the next week, because of the cooldown. 3. Poll until the state leaves `modifying`; the OS can already see the larger device. 4. Extend the partition, then grow the filesystem, and confirm with the OS's own free-space reporting. 5. Check monitoring afterwards: free space, volume queue length and latency. ## Where this fits in an interview It is usually asked as a capacity incident — "the disk is 95% full at 2am, what do you do?" — and the interviewer is listening for three things: that you know the resize is online, that you remember the in-guest step, and that you know it is one-directional so the sizing decision deserves thought. A candidate who answers "stop the instance and swap volumes" has described a decade-old procedure.

  • You ran the modification and the volume shows 500 GiB in the console, but the instance still reports the old capacity. Why?
    Because `ModifyVolume` only changes the block device AWS presents; the partition table and filesystem inside the guest are untouched and still describe the old size. You have to extend the partition to cover the new space and then grow the filesystem into it. Nothing in the EBS API can do that for you.
  • What would you do if you actually needed to shrink a volume?
    There is no shrink path in EBS. You create a new, smaller volume, attach it, copy the data at the filesystem level, verify, then swap the mount and detach the old one — which means a maintenance window. The realistic advice is to avoid the situation: over-sized storage is cheap relative to a migration.
  • Would you automate volume growth from a free-space alarm?
    Cautiously. The cooldown between modifications means a naive "add 10 GiB when 90% full" loop will be refused on its second attempt, and the in-guest filesystem extension still has to happen. If automating, grow in large steps, run the OS-side extension in the same workflow, and alert when growth was triggered so a human reviews the trend.

saying these in an interview costs you the question

  • Thinks resizing needs a stop or a detach
  • Forgets the partition and filesystem step inside the instance
  • Believes an EBS volume can be shrunk in place
  • Assumes you can modify the same volume repeatedly with no wait
  • Expects full new performance the instant the API call returns

context