skip to content

An engineer terminated an EC2 instance and the separate EBS data volume attached to it was deleted along with it. Which EC2 settings decide that outcome, and how would you stop it happening again?

level: middleimportance: should knowfreq 48%

answer

  1. a boolean per attached volume
  2. the root disk defaults the dangerous way
  3. read the live value, don't assume
  4. protection blocks the API, not everything
  5. snapshots are the actual safety net

basics

~20 s

Every entry in an instance's block device mappings carries a DeleteOnTermination flag; when it is true, terminating the instance deletes that volume. Set it to false on data volumes, enable termination protection, and keep EBS snapshots as the real safety net.

solid answer

~50 s

The deciding attribute is `DeleteOnTermination`, held per volume in the instance's block device mappings. It defaults to true for the root volume created at launch, and a volume added through a launch template or the launch wizard can easily carry it as true as well — so the honest answer is to inspect the live value with `describe-instances` rather than assume. You can flip it on a running instance with `modify-instance-attribute`. Around that, `DisableApiTermination` (termination protection) makes `TerminateInstances` fail for that instance, and `DisableApiStop` does the same for stops; both are advisory guardrails, not backups — they do not stop an Auto Scaling group from terminating the instance, a Spot interruption, or a `shutdown -h` inside the guest when `InstanceInitiatedShutdownBehavior` is set to terminate. The durable protection is snapshots on a schedule, so that a deleted volume is a restore rather than a loss.

code

bash · 7 lines
bash
aws ec2 describe-instances --instance-ids i-0123456789abcdef0 \
  --query 'Reservations[].Instances[].BlockDeviceMappings[].{dev:DeviceName,vol:Ebs.VolumeId,del:Ebs.DeleteOnTermination}'

aws ec2 modify-instance-attribute --instance-id i-0123456789abcdef0 \
  --block-device-mappings '[{"DeviceName":"/dev/sdf","Ebs":{"DeleteOnTermination":false}}]'

aws ec2 modify-instance-attribute --instance-id i-0123456789abcdef0 --disable-api-termination

go deeper

for a junior

Know that DeleteOnTermination exists as a per-volume flag, that the root volume defaults to true, and that terminating is permanent so the flag matters before you press the button.

for a middle

Be able to inspect and change the flag with describe-instances and modify-instance-attribute, and explain how termination protection and instance-initiated shutdown behaviour interact.

for a senior

Show the gaps: Auto Scaling scale-in, Spot interruptions and guest shutdowns all bypass termination protection, so your answer must end in scheduled snapshots and data that does not live on the instance.

for a principal

Own the standard rather than the setting: which classes of workload may hold durable data locally at all, how snapshot policy is enforced across accounts, and how you make disposability the default so no single flag is load-bearing.

## Where the decision actually lives An EC2 instance carries a set of **block device mappings**: one entry per attached volume, each recording the device name, the volume ID, and a boolean `DeleteOnTermination`. When the instance is terminated, EC2 walks those entries and deletes every volume whose flag is true. Nothing else is consulted — not tags, not the volume's own settings, not whether anyone is looking at it. The default that catches people is the **root volume**: created at launch with `DeleteOnTermination` set to true, on the reasoning that an instance's boot disk is part of the instance. Volumes you attach afterwards with `AttachVolume` default to false and outlive the instance as unattached volumes. Volumes created as part of the launch — in a launch template, a launch wizard, or an AMI's own block device mapping — can be either, depending on what was specified, and that ambiguity is exactly how a data volume ends up marked for deletion without anybody choosing it. So the operational rule is: never reason from the default, read the live value. ```bash aws ec2 describe-instances --instance-ids i-0123456789abcdef0 \ --query 'Reservations[].Instances[].BlockDeviceMappings[].{dev:DeviceName,vol:Ebs.VolumeId,del:Ebs.DeleteOnTermination}' ``` ## Changing it The flag is an instance attribute, so it is changed on the instance and can be changed while it runs: ```bash aws ec2 modify-instance-attribute --instance-id i-0123456789abcdef0 \ --block-device-mappings '[{"DeviceName":"/dev/sdf","Ebs":{"DeleteOnTermination":false}}]' ``` The same mechanism works the other way: on a genuinely disposable instance, marking data volumes for deletion is how you avoid accumulating orphaned volumes that quietly bill forever. ## The guardrails around termination **Termination protection** — the `DisableApiTermination` instance attribute — makes `TerminateInstances` return an error for that instance, whoever calls it, including an account administrator or the root user. You clear the attribute first, then terminate. It is a deliberate speed bump against fat fingers and against a blast-radius mistake in a script. **Stop protection** — `DisableApiStop` — is the equivalent for `StopInstances`, and matters on instances whose stop would lose instance-store data or break a licence binding. **Instance-initiated shutdown behavior** decides what a shutdown issued inside the guest means: `stop` (the default) or `terminate`. This is the classic hole in termination protection — `DisableApiTermination` blocks the *API*, and an operator typing `shutdown -h now` on a box configured with `InstanceInitiatedShutdownBehavior=terminate` is not calling the API. ```bash aws ec2 modify-instance-attribute --instance-id i-0123456789abcdef0 --disable-api-termination aws ec2 modify-instance-attribute --instance-id i-0123456789abcdef0 --disable-api-stop aws ec2 modify-instance-attribute --instance-id i-0123456789abcdef0 \ --instance-initiated-shutdown-behavior stop ``` ## What termination protection does not cover Be explicit about the gaps, because the interviewer is usually fishing for them: - An **Auto Scaling group** terminating an instance during scale-in or an instance refresh is not blocked by `DisableApiTermination`; the group has its own instance scale-in protection. - A **Spot interruption** reclaims capacity regardless of the attribute. - A **guest-initiated shutdown** with terminate behaviour bypasses it, as above. - It protects the instance, not the data. If someone clears the attribute and terminates, a `DeleteOnTermination=true` volume still goes. ## The actual answer: don't rely on any of this Flags and protections reduce accident rates; they do not make data recoverable. The engineering answer to "a volume was deleted with the instance" is that the volume should have been recoverable from a **snapshot** taken on a schedule — EBS snapshots are incremental and cheap, and they can be automated centrally so that no individual instance's configuration is what stands between you and data loss. Better still, ask why durable data was on an instance-attached volume at all: a database service, S3, or a shared file system removes the whole class of failure. The instance then becomes disposable on purpose, which is the state you want it in before an Auto Scaling group or a Spot interruption makes the decision for you.

  • Does termination protection stop an Auto Scaling group from terminating the instance?
    No. `DisableApiTermination` blocks the `TerminateInstances` API path, but an Auto Scaling group terminating an instance on scale-in or during an instance refresh is not blocked by it. If you need a specific instance spared, you use the group's own instance scale-in protection — and accept that an instance you must not lose probably should not be in an Auto Scaling group at all.
  • When would you deliberately want DeleteOnTermination set to true on a data volume?
    On genuinely disposable instances — build agents, scratch processing nodes, anything in an Auto Scaling group — where the volume holds only working data. Leaving the flag false there produces a slow accumulation of unattached volumes that keep billing and that nobody can safely identify later. The flag is a cleanup mechanism as much as a hazard.

saying these in an interview costs you the question

  • Thinks termination protection prevents data loss
  • Assumes attached data volumes are never deleted
  • Believes the flag lives on the volume rather than the instance
  • Says a guest-OS shutdown can never terminate an instance
  • Treats snapshots as unnecessary once the flag is false

context