Walk through the anatomy of a CloudWatch Logs Insights query — the pipeline of commands and the @-prefixed fields that are always available — and show how you would pull the 20 most recent lines containing ERROR out of one log group.
answer
- commands piped left to right
- @-fields exist on every event
- filter, then sort, then limit
- limit caps rows, not bytes read
- stats collapses the event stream
basics
~20 sA Logs Insights query is a pipeline of commands joined by the pipe character: fields or display picks columns, filter narrows rows, sort orders them, limit caps the output. Every event exposes @timestamp, @message, @logStream and @ingestionTime.
solid answer
~50 sCloudWatch Logs Insights queries are pipelines, not SQL. You start from the events in the log groups and time range you selected, and each command after a `|` receives the output of the one before it. `fields` (or `display`) chooses the columns, `filter` keeps matching events, `parse` extracts new fields from text, `stats` aggregates, and `sort` and `limit` shape what comes back. Every event carries system fields you can always reference: `@timestamp` (the time the producer reported), `@ingestionTime` (when CloudWatch accepted it), `@message` (the raw line), `@logStream`, and `@ptr`. JSON log lines additionally get their keys auto-discovered as dot-notation fields. So the 20 most recent errors are: `fields @timestamp, @message, @logStream | filter @message like /ERROR/ | sort @timestamp desc | limit 20`. Note that `limit` caps the rows returned — it does not reduce how much log data the query had to scan.
code
bash · 9 linesQID=$(aws logs start-query \
--log-group-name /aws/lambda/checkout \
--start-time $(( $(date +%s) - 3600 )) \
--end-time $(date +%s) \
--query-string 'fields @timestamp, @message, @logStream | filter @message like /ERROR/ | sort @timestamp desc | limit 20' \
--query queryId --output text)
sleep 5
aws logs get-query-results --query-id "$QID"go deeper
Be ready to write a working search from memory: pick columns with fields, narrow with filter and a regex, sort by @timestamp desc, then limit. Know that @message is the raw line.
Explain why the order of commands changes the result, what stats does to the fields still available downstream, and how JSON log lines get auto-discovered into dot-notation fields.
Show that you read the query statistics — recordsMatched versus recordsScanned — and that you narrow the time range and log group selection rather than adding a smaller limit when a query is slow.
Own the consequence for the platform: the shape teams log in decides whether incident queries are one filter or a fragile parse, so push structured JSON logging and a field convention rather than better queries.
## What Logs Insights actually is CloudWatch Logs Insights is a query engine that runs over log events already stored in CloudWatch Logs. You do not create an index or a table first: you pick one or more log groups and a time range in the console (or pass them to the `StartQuery` API), and the engine reads the events in that selection and runs your query over them. The query language is a **pipeline**, closer to a shell one-liner than to SQL. Commands are separated by `|` and evaluated left to right; each one receives the rows the previous one emitted. ## The system fields that are always there Even for a completely unstructured log line, every event exposes a set of `@`-prefixed fields: - `@timestamp` — the event time as reported by whatever put the log there (the agent, the SDK, the Lambda service). This is what the console's histogram and `sort @timestamp` use. - `@ingestionTime` — when CloudWatch Logs accepted the event. It differs from `@timestamp` whenever the producer buffered, retried, or had clock skew, which is exactly why the two fields exist separately. - `@message` — the raw text of the log event, unparsed. - `@logStream` — the stream inside the log group. A stream is usually one instance, task or Lambda execution environment, not one service. - `@log` — an identifier of the form `accountId:logGroupName`, which is how you tell events apart when a query spans several log groups. - `@ptr` — an opaque pointer to the stored event, used by the console's "view surrounding events" and by the `GetLogRecord` API. If the log line is JSON, Logs Insights also auto-discovers its keys as queryable fields, flattening nesting into dot notation, so `{"http":{"status":500}}` becomes the field `http.status`. There is a cap on how many fields it will discover per event, so extremely wide JSON can lose its tail. ## The commands you will actually use ``` fields @timestamp, @message, @logStream | filter @message like /ERROR/ | sort @timestamp desc | limit 20 ``` - `fields` adds columns to the output and can compute new ones (`fields @timestamp, latency / 1000 as seconds`). `display` is its blunter cousin: it *replaces* the output columns with exactly what you name, and only the last `display` in the query takes effect. - `filter` keeps events matching a boolean expression. For substring or regex matching on text use `like` (`filter @message like /ERROR/` for a regex, `filter @message like "ERROR"` for a literal substring), and `not like` for the negation. For discovered or parsed fields you can use ordinary comparisons: `filter http.status >= 500 and path = "/checkout"`. - `sort` orders rows, `asc` or `desc`. - `limit` caps the number of rows returned. - `dedup` collapses rows that share the values of the fields you name, keeping the first one — handy for "show me one example per error type". - `parse` and `stats` extract and aggregate; they are the step up from simple searching. ## Order matters, and `stats` is a wall Because it is a pipeline, position changes meaning. `sort` before `limit` gives you the top 20 by that ordering; `limit` before `sort` gives you an arbitrary 20 events that then get sorted among themselves. More importantly, `stats` **collapses** the event stream into aggregate rows: after a `stats`, the only fields that still exist are the aggregates you computed and the keys you grouped by. Referring to `@message` after a `stats` is a common beginner error. ## What `limit` does not do `limit` is an output cap, not a cost or scan control. The engine still reads the log data in the selected log groups and time range before it can decide which 20 rows to hand back. Narrowing the *time range* and the *log group selection* is what makes a query cheaper and faster; `limit` only makes the result readable. As of 2026 a single query also returns at most 10,000 log events regardless of what you ask for, so a query that is really an export needs a different tool. ## Running it outside the console Queries are asynchronous over the API: `StartQuery` returns a `queryId`, and `GetQueryResults` polls it, returning a status (`Scheduled`, `Running`, `Complete`, `Failed`) plus a `statistics` block with `recordsMatched`, `recordsScanned` and `bytesScanned`. Those three numbers are the honest feedback loop on whether your query was well targeted.
- What is the difference between `fields` and `display` in a Logs Insights query?`fields` adds columns to what is already flowing through the pipeline and can compute derived values, and several `fields` commands accumulate. `display` replaces the output columns with exactly the ones you name, and only the last `display` in the query has any effect. Use `fields` while building the query and `display` to produce a clean final table.
- Why do `@timestamp` and `@ingestionTime` sometimes differ by minutes, and which should you sort on?`@timestamp` is the time the producer stamped on the event; `@ingestionTime` is when CloudWatch Logs accepted it. Buffering in an agent, a retry after a network blip, or host clock skew opens the gap. Sort on `@timestamp` to reconstruct what happened, but check `@ingestionTime` when events seem to be missing — they may simply not have arrived yet.
- What happens to a query that references a field the events do not have?Nothing fails: the field is simply null for those events, so a `filter` on it matches nothing and a `stats` over it aggregates nothing. That is why an empty result set usually means a misspelled or never-discovered field name rather than an absence of matching logs — check by running `fields @message` alone first.
saying these in an interview costs you the question
- Writes it as SQL with SELECT and WHERE clauses
- Believes limit reduces the data scanned and billed
- Thinks @logStream identifies the service or log group
- Assumes @timestamp is when CloudWatch received the event
- Refers to @message after a stats command