skip to content

Logs Insights Querying

Interrogating logs after an incident starts. You learn the Logs Insights query language, how scanned bytes drive both latency and cost, and how to search across many log groups or accounts at once.

part ofAWSoverview, primer and where to startread it →
on this pageshow

questions

5

Walk through the anatomy of a CloudWatch Logs Insights query — the pipeline of commands and the @-prefixed fields that are always available — and show how you would pull the 20 most recent lines containing ERROR out of one log group.

level: juniorimportance: must knowfreq 60%

answer

  1. commands piped left to right
  2. @-fields exist on every event
  3. filter, then sort, then limit
  4. limit caps rows, not bytes read
  5. stats collapses the event stream

basics

~20 s

A Logs Insights query is a pipeline of commands joined by the pipe character: fields or display picks columns, filter narrows rows, sort orders them, limit caps the output. Every event exposes @timestamp, @message, @logStream and @ingestionTime.

solid answer

~50 s

CloudWatch Logs Insights queries are pipelines, not SQL. You start from the events in the log groups and time range you selected, and each command after a `|` receives the output of the one before it. `fields` (or `display`) chooses the columns, `filter` keeps matching events, `parse` extracts new fields from text, `stats` aggregates, and `sort` and `limit` shape what comes back. Every event carries system fields you can always reference: `@timestamp` (the time the producer reported), `@ingestionTime` (when CloudWatch accepted it), `@message` (the raw line), `@logStream`, and `@ptr`. JSON log lines additionally get their keys auto-discovered as dot-notation fields. So the 20 most recent errors are: `fields @timestamp, @message, @logStream | filter @message like /ERROR/ | sort @timestamp desc | limit 20`. Note that `limit` caps the rows returned — it does not reduce how much log data the query had to scan.

code

bash · 9 lines
bash
QID=$(aws logs start-query \
  --log-group-name /aws/lambda/checkout \
  --start-time $(( $(date +%s) - 3600 )) \
  --end-time $(date +%s) \
  --query-string 'fields @timestamp, @message, @logStream | filter @message like /ERROR/ | sort @timestamp desc | limit 20' \
  --query queryId --output text)

sleep 5
aws logs get-query-results --query-id "$QID"

go deeper

for a junior

Be ready to write a working search from memory: pick columns with fields, narrow with filter and a regex, sort by @timestamp desc, then limit. Know that @message is the raw line.

for a middle

Explain why the order of commands changes the result, what stats does to the fields still available downstream, and how JSON log lines get auto-discovered into dot-notation fields.

for a senior

Show that you read the query statistics — recordsMatched versus recordsScanned — and that you narrow the time range and log group selection rather than adding a smaller limit when a query is slow.

for a principal

Own the consequence for the platform: the shape teams log in decides whether incident queries are one filter or a fragile parse, so push structured JSON logging and a field convention rather than better queries.

## What Logs Insights actually is CloudWatch Logs Insights is a query engine that runs over log events already stored in CloudWatch Logs. You do not create an index or a table first: you pick one or more log groups and a time range in the console (or pass them to the `StartQuery` API), and the engine reads the events in that selection and runs your query over them. The query language is a **pipeline**, closer to a shell one-liner than to SQL. Commands are separated by `|` and evaluated left to right; each one receives the rows the previous one emitted. ## The system fields that are always there Even for a completely unstructured log line, every event exposes a set of `@`-prefixed fields: - `@timestamp` — the event time as reported by whatever put the log there (the agent, the SDK, the Lambda service). This is what the console's histogram and `sort @timestamp` use. - `@ingestionTime` — when CloudWatch Logs accepted the event. It differs from `@timestamp` whenever the producer buffered, retried, or had clock skew, which is exactly why the two fields exist separately. - `@message` — the raw text of the log event, unparsed. - `@logStream` — the stream inside the log group. A stream is usually one instance, task or Lambda execution environment, not one service. - `@log` — an identifier of the form `accountId:logGroupName`, which is how you tell events apart when a query spans several log groups. - `@ptr` — an opaque pointer to the stored event, used by the console's "view surrounding events" and by the `GetLogRecord` API. If the log line is JSON, Logs Insights also auto-discovers its keys as queryable fields, flattening nesting into dot notation, so `{"http":{"status":500}}` becomes the field `http.status`. There is a cap on how many fields it will discover per event, so extremely wide JSON can lose its tail. ## The commands you will actually use ``` fields @timestamp, @message, @logStream | filter @message like /ERROR/ | sort @timestamp desc | limit 20 ``` - `fields` adds columns to the output and can compute new ones (`fields @timestamp, latency / 1000 as seconds`). `display` is its blunter cousin: it *replaces* the output columns with exactly what you name, and only the last `display` in the query takes effect. - `filter` keeps events matching a boolean expression. For substring or regex matching on text use `like` (`filter @message like /ERROR/` for a regex, `filter @message like "ERROR"` for a literal substring), and `not like` for the negation. For discovered or parsed fields you can use ordinary comparisons: `filter http.status >= 500 and path = "/checkout"`. - `sort` orders rows, `asc` or `desc`. - `limit` caps the number of rows returned. - `dedup` collapses rows that share the values of the fields you name, keeping the first one — handy for "show me one example per error type". - `parse` and `stats` extract and aggregate; they are the step up from simple searching. ## Order matters, and `stats` is a wall Because it is a pipeline, position changes meaning. `sort` before `limit` gives you the top 20 by that ordering; `limit` before `sort` gives you an arbitrary 20 events that then get sorted among themselves. More importantly, `stats` **collapses** the event stream into aggregate rows: after a `stats`, the only fields that still exist are the aggregates you computed and the keys you grouped by. Referring to `@message` after a `stats` is a common beginner error. ## What `limit` does not do `limit` is an output cap, not a cost or scan control. The engine still reads the log data in the selected log groups and time range before it can decide which 20 rows to hand back. Narrowing the *time range* and the *log group selection* is what makes a query cheaper and faster; `limit` only makes the result readable. As of 2026 a single query also returns at most 10,000 log events regardless of what you ask for, so a query that is really an export needs a different tool. ## Running it outside the console Queries are asynchronous over the API: `StartQuery` returns a `queryId`, and `GetQueryResults` polls it, returning a status (`Scheduled`, `Running`, `Complete`, `Failed`) plus a `statistics` block with `recordsMatched`, `recordsScanned` and `bytesScanned`. Those three numbers are the honest feedback loop on whether your query was well targeted.

  • What is the difference between `fields` and `display` in a Logs Insights query?
    `fields` adds columns to what is already flowing through the pipeline and can compute derived values, and several `fields` commands accumulate. `display` replaces the output columns with exactly the ones you name, and only the last `display` in the query has any effect. Use `fields` while building the query and `display` to produce a clean final table.
  • Why do `@timestamp` and `@ingestionTime` sometimes differ by minutes, and which should you sort on?
    `@timestamp` is the time the producer stamped on the event; `@ingestionTime` is when CloudWatch Logs accepted it. Buffering in an agent, a retry after a network blip, or host clock skew opens the gap. Sort on `@timestamp` to reconstruct what happened, but check `@ingestionTime` when events seem to be missing — they may simply not have arrived yet.
  • What happens to a query that references a field the events do not have?
    Nothing fails: the field is simply null for those events, so a `filter` on it matches nothing and a `stats` over it aggregates nothing. That is why an empty result set usually means a misspelled or never-discovered field name rather than an absence of matching logs — check by running `fields @message` alone first.

saying these in an interview costs you the question

  • Writes it as SQL with SELECT and WHERE clauses
  • Believes limit reduces the data scanned and billed
  • Thinks @logStream identifies the service or log group
  • Assumes @timestamp is when CloudWatch received the event
  • Refers to @message after a stats command

context

open as a page

During an incident you need to search a dozen CloudWatch log groups belonging to different services, some of them in a second AWS account, from a single Logs Insights query. How do you do that, and what has to be in place first?

level: middleimportance: should knowfreq 40%

basics

~20 s

One Logs Insights query can span many log groups: select them individually or by name prefix, and use the @log field to tell events apart. Reaching another account first requires CloudWatch cross-account observability linking that account to your monitoring account.

open as a page

An application writes plain-text lines such as `2026-03-01T10:00:00Z level=ERROR path=/checkout latency=812ms` into a CloudWatch log group. Using CloudWatch Logs Insights, how would you turn those unstructured lines into a 95th-percentile latency per path in five-minute buckets?

level: middleimportance: should knowfreq 55%

basics

~10 s

Use parse to lift path and latency out of @message into named fields, then aggregate with stats: parse @message "path=* latency=*ms" as path, latency | stats pct(latency, 95) as p95 by path, bin(5m).

open as a page

A team says their CloudWatch Logs Insights queries take minutes to return during incidents, and the bill now shows a growing Logs Insights charge. What actually drives the cost and latency of a Logs Insights query, and how would you make the same investigation cheaper and faster?

level: seniorimportance: should knowfreq 45%

basics

~20 s

CloudWatch charges Logs Insights per gigabyte of log data scanned, and scan volume is set by the log groups and time range selected — not by how selective the filter is. Narrow the selection, not the query.

open as a page

CloudWatch Contributor Insights lets you define a rule over log groups instead of running an ad-hoc Logs Insights query. What does such a rule actually produce, and when would you build one rather than just querying?

level: seniorimportance: nice to knowfreq 25%

basics

~20 s

A Contributor Insights rule evaluates log events as they arrive and produces a continuously updated ranking of top contributors by a key you choose, plus graphable metrics you can alarm on — unlike a query, which reads history once.

open as a page