An AWS bill is dominated by Amazon CloudWatch Logs charges. Explain how CloudWatch Logs charges for data, and which levers the service itself gives you to bring the number down.
answer
- two meters, one dominates
- you pay on the way in
- retention fixes the smaller half
- some logs never needed a log group
- the class is set at creation
basics
~20 sCloudWatch Logs charges mainly for ingestion per GB, with much cheaper per-GB-month storage on compressed data. Levers: ingest less, route high-volume vended logs to S3 instead, use the Infrequent Access log class, and shorten retention — which only touches storage.
solid answer
~50 sThe bill is dominated by **ingestion**, charged per GB of uncompressed data accepted, with storage charged per GB-month on the compressed archive at a far lower rate. That ordering decides your levers. Shortening retention is the reflex answer and it is the weakest one — it only shrinks the smaller half and cannot refund bytes already ingested. The real levers are: emit less (drop debug logging in production, stop logging full request and response bodies); route high-volume AWS-vended logs such as VPC Flow Logs and ALB access logs to S3 rather than a log group, since S3 plus Athena is far cheaper for data you query occasionally; and put groups you must keep but rarely read into the **Infrequent Access log class**, which has a lower ingestion price in exchange for a reduced feature set. Note that the log class is chosen at group creation and cannot be changed afterwards.
code
bash · 13 lines# Cheaper ingestion, reduced feature set - and immutable once created
aws logs create-log-group \
--log-group-name /myapp/audit \
--log-group-class INFREQUENT_ACCESS
aws logs put-retention-policy \
--log-group-name /myapp/audit \
--retention-in-days 90
aws logs describe-log-groups \
--log-group-name-prefix /myapp \
--query 'logGroups[].[logGroupName,logGroupClass,retentionInDays,storedBytes]' \
--output tablego deeper
Know that CloudWatch Logs charges both for data ingested and for data stored, and that the amount you log is what drives the bill.
Explain the two meters and their relative size — per-GB ingestion on uncompressed data versus cheaper per-GB-month storage on compressed data — and why retention therefore only affects the smaller half.
Investigate rather than guess: find the few log groups producing most of the volume, cut what is logged, route high-volume vended logs to S3, and use the Infrequent Access class knowing it is fixed at creation and loses features.
Own the policy — which telemetry is allowed into CloudWatch Logs at all, tagging so ingestion is attributable to a team, and the standing trade between an operable log platform and a cheaper archive that nobody can search during an incident.
## The two meters CloudWatch Logs bills on: 1. **Ingestion (collection)** — per GB of data accepted into log groups. AWS-vended logs (VPC Flow Logs, Route 53 Resolver logs, API Gateway access logs and similar) are metered on a separate vended-logs tier with volume discounts, but they are still charged on the way in. 2. **Storage (archival)** — per GB-month, measured on the compressed stored data. For virtually every real workload, ingestion is the larger number by an order of magnitude, because storage is both discounted and compressed. Interactive querying is metered separately again on data scanned, and any metrics you derive from logs are billed as custom metrics. Internalise the ordering, because it determines which levers matter: **you pay for a log line once, permanently, the moment it is ingested.** Nothing you do afterwards refunds that. ## Lever 1 — ingest less (the only one that moves the big number) Unglamorous and dominant: - Turn off debug-level logging in production. A single library left at DEBUG can be most of a bill. - Stop logging full request and response payloads; log identifiers and outcomes. - Sample high-frequency success paths and keep every failure. - Deduplicate — health-check access lines, per-item loop logging, retry storms that log every attempt. - Check what is being logged twice: an application writing to stdout *and* to a file the agent tails produces two copies of the same bytes in two log groups. ## Lever 2 — do not send high-volume telemetry through CloudWatch Logs at all Several AWS services let you choose the destination for their logs, and CloudWatch Logs is not always the right one. **VPC Flow Logs** and **ALB access logs** are the classic examples: enormous volume, read rarely, and usually read in bulk rather than interactively. Delivered to S3 they cost object storage prices and can be queried with Athena when actually needed. Route them into a log group and you pay vended-logs ingestion for every record forever. The decision rule is simple: does a human tail this, or correlate it with application logs in an incident? If not, it belongs in S3. ## Lever 3 — the Infrequent Access log class CloudWatch Logs offers two log classes, `STANDARD` and `INFREQUENT_ACCESS`, selected with `--log-group-class` at creation: ```bash aws logs create-log-group \ --log-group-name /myapp/audit \ --log-group-class INFREQUENT_ACCESS ``` Infrequent Access has a materially lower per-GB ingestion price (about half at launch) in exchange for a reduced capability set — notably no Live Tail and no metric extraction via metric filters, among other omissions. It fits data you are obliged to retain and occasionally search but never operate against: audit trails, compliance logs, verbose debug streams kept just in case. Two practical constraints. First, **the class is fixed at creation** — you cannot flip an existing group, so adopting it means creating new groups and cutting over. Second, check the current capability matrix before committing, because AWS has been adding features to the class over time and an old blog post is not authoritative. ## Lever 4 — retention, in its proper place Setting retention on groups that default to *never expire* is worth doing, and on an old account with years of accumulated data it can be a visible saving. But it addresses storage, the smaller meter. Treat it as hygiene rather than as the answer to "the bill is too high". Where retention becomes a real lever is in combination with export: keep a short retention on the log group and stream everything out with a subscription filter to Amazon Data Firehose into S3, where lifecycle rules move objects to Glacier tiers. You still pay ingestion once, but you stop paying CloudWatch storage for a seven-year compliance window at CloudWatch prices. Be honest in an interview that this adds a second pipeline to operate — it is a trade, not a free win. ## Making the number attributable You cannot fix what you cannot attribute. Tag log groups so their cost lands against a team, and look at per-log-group ingestion volume — in almost every investigation, a small number of groups account for most of the bill, and one of them is usually a debug flag someone left on. Start there rather than with a blanket retention policy that annoys everyone and saves little. ## The answer that lands Lead with "ingestion dominates, so the lever is volume", then name vended-log routing to S3, the Infrequent Access class with its creation-time constraint, and retention as the storage-side cleanup. Leading with retention is the tell that someone has read the console but not a bill.
- Why is shortening retention a weak answer to a high CloudWatch Logs bill?Because it only touches storage, which is charged per GB-month on compressed data and is typically a fraction of the total. Ingestion is charged once per GB on the way in and is never refunded, so a group with a huge daily volume costs almost the same whether you keep it for a week or a year. Volume reduction is the lever.
- How would you find which log groups are actually driving the cost?Look at per-log-group ingested volume rather than at the invoice total — CloudWatch publishes IncomingBytes per log group, and describe-log-groups reports storedBytes. Almost every investigation ends at a handful of groups, and one of them is usually a service left at debug level or logging full payloads. Tag groups by team so the finding has an owner.
- What do you give up by choosing the Infrequent Access log class?Capability, in exchange for a lower ingestion price. Notably you lose Live Tail and metric extraction via metric filters, so it is unsuitable for anything you operate against or alarm on. It fits retain-and-rarely-read data such as audit trails. And because the class is fixed at creation, adopting it means new groups and a cutover, not a setting change.
- When is streaming logs out to S3 not worth it?When the volume is modest or the retention window is short — you still pay CloudWatch ingestion for every byte, and you have added a Firehose pipeline and a bucket to operate for a saving measured against already-cheap compressed storage. The trade pays off for large volumes kept for years, not for a 30-day application log.
saying these in an interview costs you the question
- Naming retention as the main lever on a CloudWatch Logs bill
- Assuming storage costs more than ingestion
- Thinking the log class can be switched on an existing group
- Believing streaming logs to S3 avoids the ingestion charge
- Sending VPC Flow Logs to CloudWatch Logs by default