skip to content

A team says their CloudWatch Logs Insights queries take minutes to return during incidents, and the bill now shows a growing Logs Insights charge. What actually drives the cost and latency of a Logs Insights query, and how would you make the same investigation cheaper and faster?

level: seniorimportance: should knowfreq 45%

answer

  1. billed per gigabyte scanned
  2. selection decides the scan, not the filter
  3. time range is the biggest lever
  4. limit and filter shape output only
  5. recurring questions belong in metrics

basics

~20 s

CloudWatch charges Logs Insights per gigabyte of log data scanned, and scan volume is set by the log groups and time range selected — not by how selective the filter is. Narrow the selection, not the query.

solid answer

~60 s

A Logs Insights query is billed and paced by **bytes scanned**, and by default the engine reads every log event in the log groups and time range you selected before your `filter` gets a say. That is why adding a tighter filter or a smaller `limit` changes nothing: those shape the result, not the read. The levers that actually work are all about the selection — shorten the time range (start with fifteen minutes around the incident and widen only if needed), select fewer and more specific log groups instead of a broad prefix, and split noisy debug output into its own log group at emit time so it is not dragged into every query. If you have a high-selectivity field you filter on constantly, a field index policy lets the engine skip data that cannot match. And any question you ask on a schedule should stop being a log scan at all — publish it as a metric so the recurring answer costs nothing to read. Note this charge is separate from log ingestion and storage.

go deeper

for a junior

Know that Logs Insights charges by how much log data the query reads, and that shrinking the time range is what makes a query cheaper — not a smaller limit.

for a middle

Explain why the filter cannot reduce scanned bytes, and read recordsScanned against recordsMatched to judge whether a query was aimed properly.

for a senior

Demonstrate incident discipline: narrow window first, iterate the query there, widen once. Then name the structural fixes — log group separation, index policies, promoting recurring questions to metrics.

for a principal

Own the tradeoff between retaining everything cheaply and being able to interrogate it affordably; set the log group layout, retention and metric conventions that keep incident queries small by construction.

## The billing unit is the scan, not the answer CloudWatch Logs Insights is priced per gigabyte of **log data scanned** by a query. That is a different line from what you pay to ingest and to store the logs; a log group you never query still costs money, and a log group you query fifty times a day costs money fifty extra times. Latency follows the same quantity: a query that has to read hundreds of gigabytes takes minutes because it is reading hundreds of gigabytes. Crucially, **the amount scanned is decided before your query logic runs**. You select log groups and a time range; the engine reads the events in that box. A `filter` that matches three events out of forty million still had to look at forty million, and a `limit 20` truncates the result after the work is done. This is the mental model most people get wrong, and it is why "my query is expensive so I made the filter stricter" never helps. The honest feedback loop is in the query statistics: `GetQueryResults` returns `recordsScanned`, `recordsMatched` and `bytesScanned`, and the console shows the same numbers under the results. A query with forty million scanned and twelve matched is a badly targeted query even if it returned the right answer. ## The levers that actually reduce scanned bytes **Time range.** This is the biggest and cheapest lever. Scan volume is roughly linear in the window. During an incident, start at fifteen or thirty minutes around the first alarm and widen only when you must; do not open a query on "last 3 days" out of habit. The console's event histogram is the fast way to find where the interesting window is before you run the expensive query. **Log group selection.** Every log group you add multiplies the read. Selecting a prefix such as `/aws/lambda/` because it is convenient can pull in dozens of unrelated functions. Name the two or three log groups that could plausibly hold the answer. **Log group design, upstream.** If a service writes debug chatter and business events into the same log group, every query for a business event pays for the chatter. Separating them at emit time — different log groups, or a lower log level in production — reduces the scan on every future query. This is the structural fix and it is usually the one that pays. **Field index policies.** CloudWatch Logs lets you declare index policies on selected fields (`PutIndexPolicy`); when a query filters on an indexed field with an equality match, the engine can skip log data that cannot contain a match, cutting both time and scanned bytes. It only helps high-selectivity fields you filter on repeatedly — a request id or a tenant id, not a log level with four values — and only for queries written to filter on that field directly. **Stop querying for recurring questions.** Anything you ask on a schedule, or that backs a dashboard or an alert, should not be an ad-hoc log scan. Publish the number as a metric from the application, and the recurring question is answered by reading a metric instead of rescanning gigabytes of text. Ad-hoc querying is for the questions you did not anticipate. ## Things people try that do not help - **A smaller `limit`.** Output cap only. - **A stricter `filter` or a cleverer regex.** Applied after the read; it may make the query marginally faster to evaluate but it does not change what was read. - **Sorting differently, or dropping columns with `display`.** Cosmetic to the scan. - **Re-running the same broad query five times while iterating on the regex.** This is the real incident-time cost driver: each iteration pays the full scan. Iterate against a narrow window first, get the query right, then widen it once. ## Operational judgment There is also a concurrency ceiling on how many Logs Insights queries an account can run at once, so a room full of responders each running a three-day query during an incident degrades everyone. The team-level habit worth teaching is: narrow first, save the query once it works, and promote anything you have now run three times into a metric or a Contributor Insights rule. Cost and incident latency improve together, because they are the same number.

  • Two engineers run the same query, one over 30 minutes and one over 7 days. How do the costs compare?
    Roughly in proportion to the log volume in each window, so the seven-day query costs on the order of three hundred times as much and takes correspondingly longer — for the same log groups. Scan volume is essentially linear in the time range, which is why narrowing the window is the first and largest lever available during an incident.
  • How do you tell whether a query was well targeted after it returns?
    Compare `recordsScanned` with `recordsMatched` in the query statistics, and look at `bytesScanned`. Millions scanned for a handful matched means the selection was too broad — the fix is a shorter range or fewer log groups, not a stricter filter. Those three numbers are also the right thing to quote when justifying a change to log group layout.
  • When does a field index policy fail to help?
    When the field is low-cardinality, such as a log level, because almost no data can be skipped; when the query filters on the raw `@message` text rather than the indexed field; and for data already ingested before the policy existed. It is a targeted optimisation for a hot equality filter, not a general index over the log group.

saying these in an interview costs you the question

  • Thinks a stricter filter reduces the bytes billed
  • Believes limit stops the engine reading early
  • Confuses Logs Insights scan cost with log ingestion cost
  • Opens every incident query on a multi-day range
  • Runs a dashboard's recurring question as a log query

context