skip to content

You inherit an AWS account whose monthly bill keeps climbing although no new workload has shipped for a year. Which classes of resource keep charging after whatever needed them is gone, and how would you find them?

level: middleimportance: should knowfreq 45%

answer

  1. storage outlives the compute
  2. deleting the parent leaves the residue
  3. hourly charges do not care about traffic
  4. state names the waste: available, unassociated
  5. inventory finds it, the bill does not

basics

~20 s

Storage and reserved capacity outlive the compute that created them: unattached EBS volumes, accumulating snapshots and AMIs, public IPv4 addresses now billed hourly whether used or not, idle load balancers, and non-production environments running around the clock.

solid answer

~50 s

The pattern is that **provisioned capacity keeps billing while nothing uses it**, and deleting the obvious resource rarely deletes what it left behind. The usual suspects: EBS volumes left in the `available` state after their instance was terminated; snapshots and the AMIs backed by them, which survive the volume's deletion; public IPv4 addresses, which since February 2024 are billed hourly whether attached or idle; load balancers and NAT Gateways with no traffic still paying their hourly charge; provisioned concurrency or provisioned IOPS left enabled after a load test; and dev or staging environments running nights and weekends. To find them you need a resource inventory, not a cost report — the bill tells you what is expensive, not what is unused. Enumerate by state (`describe-volumes` filtered on `available`, addresses with no association), then cross-check candidates against a CloudWatch metric that would be non-zero if anything were using them.

code

bash · 5 lines
bash
# EBS volumes billing while attached to nothing
aws ec2 describe-volumes \
  --filters Name=status,Values=available \
  --query 'Volumes[].{Id:VolumeId,GiB:Size,Type:VolumeType,AZ:AvailabilityZone,Created:CreateTime}' \
  --output table

go deeper

for a junior

Know that storage keeps billing even when nothing is attached to it, and that a resource you cannot see in the console is still on the invoice until it is deleted.

for a middle

Be able to name the classes — available volumes, orphaned snapshots and AMIs, hourly-billed public IPv4 addresses, idle load balancers — and the state filters that enumerate each one.

for a senior

Show the safe method: inventory by state, confirm with a usage metric over a real window, delete reversibly in batches, and follow up with lifecycle defaults so the waste does not return.

for a principal

Frame it as an ownership problem rather than a cleanup problem — the durable fix is defaults and accountable ownership at creation time, not a quarterly sweep somebody has to volunteer for.

## Why waste accumulates invisibly Every AWS account drifts toward waste for the same structural reason: **creating a resource is a one-line action with an owner, and deleting it is nobody's job.** Compute is usually noticed because it is the largest line, but the residue compute leaves behind is small per item, permanent, and never appears in a deploy diff. A hundred forgotten 100 GiB volumes cost real money forever and appear in no architecture diagram. ## The classes worth sweeping **Detached block storage.** Terminating an EC2 instance does not necessarily delete its volumes; any volume without `DeleteOnTermination` set survives in the `available` state and bills the full provisioned GB-month rate. It serves no I/O and nothing references it. **Snapshots and AMIs.** Deleting a volume does not delete its snapshots. Snapshots are incremental, so a long chain is cheaper than its nominal total, but chains from an application decommissioned two years ago are pure waste. Deregistering an AMI does not delete the snapshots backing it either — a very common half-cleanup that leaves the storage behind. **Public IPv4 addresses.** Historically only *unassociated* Elastic IPs were charged. Since 1 February 2024 AWS bills **every** public IPv4 address per hour, associated or not. That turned a small hygiene issue into a real line item on accounts holding dozens of addresses for instances that no longer need public reachability. **Idle hourly infrastructure.** Load balancers, NAT Gateways, interface VPC endpoints, Transit Gateway attachments and managed NAT-like services all carry standing hourly charges. A load balancer left in front of a decommissioned service serves zero requests and bills continuously. **Provisioned capacity left switched on.** Provisioned IOPS on a volume sized for a migration that finished; Lambda provisioned concurrency enabled for a launch that is over; over-provisioned throughput on a data store. These are invisible because the resource itself is still legitimately in use — only the *setting* is waste. **Always-on non-production.** Dev and test environments running 168 hours a week to serve 40 hours of use is often the single largest saving in a neglected account, and the least controversial. **Volume type left on an old generation.** Volumes still on `gp2` can usually move to `gp3`, which is priced roughly 20% lower per GB and provisions a baseline of IOPS and throughput independently of size rather than tying performance to capacity. The migration is in place and does not require detaching. ## How to find them The important distinction: **cost tooling shows you what is expensive; only a resource inventory shows you what is unused.** A forgotten volume is not anomalous spend — it has been billing steadily for a year. Work in two passes: 1. **Enumerate by state.** Most waste has a state that names it: volumes in `available`, addresses with no association, images with no running instance, log groups with no recent writes, load balancers with no registered healthy targets. The CLI answers most of this directly. 2. **Confirm with a usage signal.** Before deleting anything, check a CloudWatch metric that would be non-zero if the resource were live — request counts on a load balancer, processed bytes on a NAT Gateway, invocations on a function. Absence of traffic over 30 days is a much better delete signal than absence of a tag. Then make deletion safe rather than brave: snapshot a volume before deleting it, delete in a batch with a rollback window, and start with the resources whose loss is recoverable. ## Keeping it from coming back One-off sweeps decay. What holds is making ownership legible — an owner tag applied at creation so an unclaimed resource is visibly unclaimed — plus lifecycle rules that expire snapshots automatically, retention set on log groups at creation rather than left on the default, and a scheduled stop for non-production outside working hours. The sweep is the fix; the defaults are the cure.

  • How do you make a delete sweep safe when you cannot prove a resource is unused?
    Bias toward reversible steps. Snapshot a volume before deleting it, detach or disable before destroying, and stop instances for a cooling-off period before termination. Batch changes with an announced window so an owner can object, and start with resource classes whose loss is recoverable from backup rather than the ones that hold the only copy of something.
  • Why is moving gp2 volumes to gp3 usually an easy win?
    gp3 is priced roughly 20% lower per GB and provisions baseline IOPS and throughput independently of volume size, whereas gp2 ties performance to capacity — so teams often over-size gp2 volumes purely to buy IOPS. The change is in place and does not require detaching, which makes it one of the least disruptive levers available.
  • What keeps this waste from re-accumulating after the sweep?
    Defaults, not discipline. Set snapshot expiry with a lifecycle policy, set log retention at creation instead of leaving the default, require an owner tag so unclaimed resources are visible, and schedule non-production environments to stop outside working hours. Anything that depends on someone remembering to clean up will drift again.

saying these in an interview costs you the question

  • Assumes terminating an instance removes its volumes and snapshots
  • Thinks an idle load balancer costs nothing without traffic
  • Believes only unassociated Elastic IPs are billed
  • Deletes candidates without checking a usage metric first
  • Treats a one-off sweep as a permanent fix

context