skip to content

Cost Management & Optimization

Making an AWS bill explainable and then smaller. This branch covers attributing spend to teams and workloads, analysing and forecasting it, choosing a purchasing model, and the concrete levers that cut waste.

part ofAWSoverview, primer and where to startread it →
on this pageshow

explore

questions

21

Your AWS bill is up sharply month over month and nobody knows why. Walk through how you would use AWS Cost Explorer — granularity, Group by and filters — to narrow the increase down to a specific cause.

level: juniorimportance: must knowfreq 70%

answer

  1. narrow by when, then what, then why
  2. daily granularity before monthly totals
  3. group by service, then usage type
  4. billing data is roughly a day behind
  5. usage type names the actual billed thing

basics

~20 s

Set AWS Cost Explorer to Daily granularity to find the day spend jumped, Group by Service to name the culprit, then filter to that service and regroup by Usage Type, Region or Linked Account until the charge line is identified.

solid answer

~50 s

I work it as three narrowing passes. First **when**: switch granularity from Monthly to Daily so the increase shows up as a step or a ramp on a specific date — that alone often matches it to a deploy or a launch. Second **what**: keep the daily view and Group by Service, so I can see which service's line moved rather than reading a single total. Third **why**: filter down to that one service and re-group by Usage Type, which is the dimension that names the actual billed thing (`BoxUsage`, `DataTransfer-Out-Bytes`, `TimedStorage-ByteHrs`), then by Region, Instance Type or API Operation as needed. In an organization I also group by Linked Account early, to see whose account moved. Two caveats: Cost Explorer data refreshes at most about daily, so today is incomplete, and it aggregates line items — it will not name a resource id.

code

bash · 5 lines
bash
aws ce get-cost-and-usage \
  --time-period Start=2026-07-01,End=2026-08-01 \
  --granularity DAILY \
  --metrics UnblendedCost \
  --group-by Type=DIMENSION,Key=SERVICE

go deeper

for a junior

Be ready to name the controls out loud: granularity, Group by, filter. Say you would switch to Daily, group by Service, then drill into Usage Type — that sequence alone is what the question is testing.

for a middle

Explain what each dimension actually means — usage type versus API operation versus charge type — and why the data is roughly a day behind. Mention that hourly and resource-level detail are opt-in rather than assuming they are available.

for a senior

Show diagnostic judgment: read the shape of the daily curve (step, ramp, sawtooth) as evidence about the cause, separate rate changes from quantity changes, and rule out credits and month length before escalating to a team.

for a principal

Own the question of how a spend investigation should not depend on one person's console skill: what standing views, exports and per-account ownership exist so the first pass is already done when someone asks, and what that costs to maintain.

## What Cost Explorer actually shows AWS Cost Explorer is a reporting front end over your billing data — the same line items that make up the invoice, aggregated and charted. Two consequences follow immediately, and candidates who miss them chase ghosts. It is **not real time**. The billing pipeline refreshes the data at most about once a day, so the current day is always partial and even yesterday can still settle. If you look at a spike an hour after it happened, you will not see it. For minute-level signal you are in CloudWatch metrics territory, not billing. It is **aggregated**. Every point is a sum of billing line items grouped by whatever dimensions you chose. Cost Explorer answers "which kind of charge grew", not "which specific bucket or instance did it". As of 2026 it keeps roughly the last year of history for charting and can forecast forward, and it offers Monthly and Daily granularity by default; **Hourly granularity, and resource-level detail, are an opt-in setting that carries its own charge and a short retention window**. Do not assume they are on. ## The three narrowing passes **Pass 1 — when.** Switch granularity to Daily and widen the range to cover both months. The shape tells you the class of problem. A vertical step on one date usually means something was switched on: a new environment, a new service, a scale-out. A gradual ramp usually means accumulation — storage growing, log retention set to never expire, snapshots piling up. A sawtooth that only appears on weekdays points at scheduled or human-driven work. **Pass 2 — what.** Keep Daily and set **Group by: Service**. A single total is useless; the stacked-by-service chart shows exactly which band grew. In a multi-account organization, Group by **Linked Account** at the same time is often faster — the increase frequently belongs to one team's account, and that tells you who to talk to. (Attributing cost by team via tags and Cost Categories is a separate discipline, handled elsewhere.) **Pass 3 — why.** Now add a **filter** on the service you identified, and re-group. The dimension that usually cracks it is **Usage Type**, because usage types are the literal billed things and their names are self-describing: `USE1-BoxUsage:m5.xlarge` is on-demand compute hours in us-east-1, `TimedStorage-ByteHrs` is stored bytes over time, `DataTransfer-Out-Bytes` is egress, `Requests-Tier1` is request counts. Other dimensions worth knowing: **API Operation** (which call is being made — invaluable for S3 and other request-priced services), **Region**, **Instance Type**, **Availability Zone**, **Purchase Option** (On-Demand versus Spot versus reserved), and **Charge Type**, which separates plain usage from tax, credits, refunds and reservation fees. Filter and Group by are complementary: filter removes rows from consideration, Group by splits what remains into series. Filtering to one service and grouping by usage type is the workhorse combination. ## Reading the result honestly A few things routinely mislead: - **Month length.** February against March is a three-day handicap on anything hourly. Compare daily run rates, not month totals. - **Credits and refunds.** A credit expiring makes cost "rise" with no change in usage at all. Check the Charge Type dimension and the include/exclude settings for credits, refunds, taxes and support before declaring an incident. - **Which cost metric is selected.** The default is unblended cost; a month containing an upfront commitment payment will look alarming under that metric and calm under amortized. - **Rate versus quantity.** Cost is rate times usage. Cost Explorer can chart **Usage quantity** as well as cost — if quantity is flat and cost moved, the cause is a pricing or discount change, not your workload. ## Where the drill-down stops When the usage type is identified but you still need the individual resource — which of four hundred buckets, which volume — Cost Explorer has run out of resolution unless resource-level granularity is enabled. Per-resource attribution at that point comes from the detailed billing exports queried directly, which is a different tool and a different leaf. One last practical note: the Cost Explorer **API** is billed per request, so a script polling `GetCostAndUsage` on a tight loop is itself a cost line. Cache, and query with the granularity you actually need.

  • You have found the day the cost stepped up, but the daily chart still shows no obvious service. What next?
    Two moves. Group by Charge Type to check whether the step is usage at all — an expiring credit or a tax change moves cost with flat usage. If it is genuinely usage, chart Usage quantity beside cost: a flat quantity with rising cost points at a rate or discount change, such as a commitment expiring, rather than at anything your workload did.
  • Why would you group by API Operation rather than Usage Type?
    For request-priced services, the usage type may only say `Requests-Tier1` while the API Operation dimension names the specific call — `GetObject`, `PutObject`, `ListBucket`. That is what turns "request charges grew" into "something is listing this bucket in a loop", which is an actionable finding.
  • Cost Explorer shows almost nothing for today. Is something broken?
    No — that is expected. Billing data refreshes at most about once a day, so the current day is partial and recent hours are missing entirely. Never diagnose a live incident from Cost Explorer; use it the next day, and use CloudWatch metrics for anything that needs to be seen within minutes.

saying these in an interview costs you the question

  • Expecting Cost Explorer to update in real time
  • Comparing month totals without adjusting for month length
  • Reading one total instead of grouping by service
  • Assuming Cost Explorer names the individual resource
  • Ignoring credits and refunds as a cause of an apparent rise

context

open as a page

AWS bills EC2 compute at On-Demand rates by default. What are the main ways to pay less for the same compute, and what does each one ask of you in return?

level: juniorimportance: must knowfreq 76%

basics

~20 s

Three purchase axes exist. On-Demand pays the list rate for total flexibility. Commitments — Savings Plans or Reserved Instances — trade a one- or three-year pledge for a large discount. Spot buys spare capacity cheaply, but AWS can reclaim it.

open as a page

Your EC2 instances and S3 buckets already carry a `CostCenter` tag, but AWS billing reports still show that spend as untagged. Explain what a cost allocation tag is, how user-defined tags differ from AWS-generated ones, and what has to happen before a tag can group costs.

level: middleimportance: must knowfreq 60%

basics

~20 s

Tagging a resource does not change billing by itself. A tag key only groups cost after it is activated as a cost allocation tag in the Billing console of the management (payer) account, which takes up to 24 hours and applies to usage from then on.

open as a page

A team asks you to "cap" their AWS spend at a fixed amount each month. Explain what AWS Budgets can and cannot do for them, including actual versus forecasted alert thresholds and what budget actions add.

level: middleimportance: must knowfreq 62%

basics

~20 s

AWS Budgets notifies, it does not cap. A cost budget alerts on actual or forecasted thresholds; only an attached budget action — applying a restrictive IAM policy or SCP, or stopping EC2 and RDS instances — changes anything on its own.

open as a page

On an AWS bill, which network traffic is free and which shows up as a data-transfer charge? Cover traffic between two instances in one Availability Zone, traffic between Availability Zones in the same Region, and traffic out to the internet.

level: middleimportance: must knowfreq 65%

basics

~20 s

Inbound internet traffic is free; outbound is charged per gigabyte. Traffic crossing Availability Zones inside a Region is billed in both directions. Same-AZ traffic is free over private IPv4 but charged when it uses public or Elastic IP addresses.

open as a page

Compare an AWS Compute Savings Plan with an EC2 Instance Savings Plan: what does each one commit you to, and what flexibility do you give up for the deeper discount?

level: middleimportance: must knowfreq 64%

basics

~20 s

Both commit a fixed dollar-per-hour spend for one or three years. A Compute Savings Plan applies to EC2 in any family or region and also to Fargate and Lambda. An EC2 Instance Savings Plan is locked to one instance family in one region and discounts more deeply.

open as a page

AWS Cost Explorer lets you chart cost as Unblended, Blended or Amortized. What does each metric mean, and which one do you use to explain a month that contained a large upfront Savings Plan or Reserved Instance payment?

level: middleimportance: should knowfreq 52%

basics

~20 s

Unblended is the rate an account was actually charged as usage occurred; blended averages rates across a consolidated billing family; amortized spreads upfront Reserved Instance and Savings Plan fees over the commitment term. Use amortized for the upfront month.

open as a page

An AWS Lambda function is billed in GB-seconds. A colleague proposes cutting cost by lowering its memory setting from 1024 MB to 512 MB. Explain when that actually lowers the bill and when it raises it.

level: middleimportance: should knowfreq 48%

basics

~20 s

Lambda bills memory multiplied by duration, and CPU is allocated in proportion to memory. Halving memory halves the per-millisecond price but can more than double the runtime of CPU-bound work, raising total cost. It only saves money for functions that mostly wait on I/O.

open as a page

You inherit an AWS account whose monthly bill keeps climbing although no new workload has shipped for a year. Which classes of resource keep charging after whatever needed them is gone, and how would you find them?

level: middleimportance: should knowfreq 45%

basics

~20 s

Storage and reserved capacity outlive the compute that created them: unattached EBS volumes, accumulating snapshots and AMIs, public IPv4 addresses now billed hourly whether used or not, idle load balancers, and non-production environments running around the clock.

open as a page

Savings Plans have largely replaced Reserved Instances for new AWS EC2 commitments. What can a Reserved Instance still do that a Savings Plan cannot, and when would you buy a standard rather than a convertible RI?

level: middleimportance: should knowfreq 52%

basics

~20 s

Reserved Instances can reserve capacity when scoped to an Availability Zone, can be exchanged if convertible, and can be sold on the Reserved Instance Marketplace if standard. Savings Plans do none of these — and services such as RDS and ElastiCache still only offer reserved instances.

open as a page

Finance wants hourly cost per team tag, broken down by usage type, for the last twelve months — more detail than the Billing console will show. Explain what the AWS Cost and Usage Report gives you that the console does not, and how you would query it.

level: seniorimportance: should knowfreq 40%

basics

~20 s

The Cost and Usage Report delivers the raw billing line items to an S3 bucket you own, at hourly granularity with activated tags and optionally resource IDs as columns. You query it with Athena over the S3 data, which gives detail and joins the console cannot express.

open as a page

Teams in your AWS Organization keep launching resources with no `CostCenter` tag, or spelling it `costcenter`. Explain what an AWS Organizations tag policy actually enforces, and how you would prevent untagged resources from being created at all.

level: seniorimportance: should knowfreq 45%

basics

~20 s

An Organizations tag policy standardises tag keys and allowed values and reports non-compliance; with enforcement enabled for named resource types it blocks non-compliant tagging operations. It never blocks creating an untagged resource — that needs a deny policy on the aws:RequestTag condition key.

open as a page

A monthly AWS cost budget only tells you once spend is already high. How does AWS Cost Anomaly Detection differ — what a monitor watches, how alert thresholds are configured, and what it still cannot catch?

level: seniorimportance: should knowfreq 42%

basics

~20 s

AWS Cost Anomaly Detection learns each monitored dimension's normal spend pattern and alerts on unexpected deviation, catching a spike days before a monthly threshold. It cannot see spend that has always been high, or a slow creep that resembles growth.

open as a page

The largest line on a VPC's monthly bill is 'NAT Gateway data processing'. The workload is a fleet of containers in private subnets that pull images from Amazon ECR and read and write objects in Amazon S3 all day. Explain why that charge is so large and what you would change to reduce it.

level: seniorimportance: should knowfreq 55%

basics

~20 s

A NAT Gateway bills an hourly rate plus a per-GB data-processing fee on everything it forwards — including in-Region AWS traffic. Route S3 and DynamoDB through free gateway VPC endpoints and other services through interface endpoints so that bulk traffic never touches the NAT.

open as a page

Finance reports that your AWS Savings Plans utilization is 100% but coverage is around 40%. What do those two numbers measure, and what does that combination tell you to do next?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Utilization is the share of your committed hourly spend that eligible usage actually consumed; coverage is the share of eligible compute usage that a commitment discounted. Full utilization with low coverage means the commitment is sound but too small — most compute is still billing at On-Demand rates.

open as a page

You own FinOps for an AWS Organization of around 120 accounts and leadership wants each product team charged for what it uses. How would you design the allocation model, and what would you do about spend that no tag can attribute?

level: principalimportance: should knowfreq 36%

basics

~20 s

Make the account boundary carry most of the allocation, since all spend in an account belongs to it without tagging discipline, then use a small mandatory tag set inside accounts. Split genuinely shared cost by rule, and start with showback before charging anyone.

open as a page

Your company wants a multi-year AWS compute commitment, but the fleet is mid-migration to containers and part of it is moving to a different instance architecture. How do you decide what to commit to, and for how long?

level: principalimportance: should knowfreq 34%

basics

~20 s

Commit only to the part of the baseline that survives every plausible roadmap outcome, prefer Compute Savings Plans because they follow workloads across families and into Fargate and Lambda, ladder shorter terms over the uncertain portion, and buy in tranches rather than one irreversible purchase.

open as a page

What are AWS Cost Categories in the Billing and Cost Management console, how do they differ from cost allocation tags, and what problem do their split charge rules solve?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

A Cost Category is a billing dimension you define with rules over accounts, tags, services, regions and charge types, rather than one that resources carry. It groups spend that tags cannot reach, and its split charge rules distribute shared costs across the teams that caused them.

open as a page

AWS Cost Explorer provides both a utilization report and a coverage report for Savings Plans and Reserved Instances. What is the difference between the two, and what does a low value on each one tell you?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Utilization is the share of a purchased commitment that was actually used; coverage is the share of eligible on-demand usage that a commitment discounted. Low utilization means you are paying for unused commitment; low coverage means discountable usage is still billed at On-Demand rates.

open as a page

A manager wants to buy a three-year AWS Savings Plan immediately to cut the EC2 bill. Why should rightsizing and a Graviton evaluation happen first, and what does AWS Compute Optimizer contribute?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

A commitment discounts the fleet you have, so buying first locks in three years of existing waste. Rightsizing and moving eligible workloads to cheaper Graviton instance types lower the baseline first; AWS Compute Optimizer supplies the per-resource evidence for those changes.

open as a page

Cross-AZ data transfer between chatty internal services is the largest line on your AWS bill. An engineer proposes making every service call only same-Availability-Zone instances of its dependencies. How would you evaluate that proposal?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Zone-local routing removes a real per-gigabyte charge but trades away the load spreading and headroom that multi-AZ deployment buys. Quantify the saving first, then require per-zone capacity headroom and an automatic fallback to remote zones before accepting it.

open as a page