skip to content

An application writes plain-text lines such as `2026-03-01T10:00:00Z level=ERROR path=/checkout latency=812ms` into a CloudWatch log group. Using CloudWatch Logs Insights, how would you turn those unstructured lines into a 95th-percentile latency per path in five-minute buckets?

level: middleimportance: should knowfreq 55%

answer

  1. glob stars capture in order
  2. regex form uses named groups
  3. stats collapses to aggregate rows
  4. bin() makes the time bucket
  5. non-matching lines survive with nulls

basics

~10 s

Use parse to lift path and latency out of @message into named fields, then aggregate with stats: parse @message "path=* latency=*ms" as path, latency | stats pct(latency, 95) as p95 by path, bin(5m).

solid answer

~50 s

Two commands do the work. `parse` extracts ephemeral fields from the raw text — either with a glob pattern where each `*` is captured in order, or with a regex using named capture groups when the text is messier. Then `stats` aggregates: `pct(latency, 95)` gives the percentile, and grouping keys go after `by`. To get time buckets you group by `bin(5m)`, which rounds `@timestamp` down to a five-minute boundary, alongside the ordinary key. The whole thing reads: `parse @message "path=* latency=*ms" as path, latency | stats pct(latency, 95) as p95, count(*) as n by path, bin(5m) | sort p95 desc`. Two things bite people: events that do not match the parse pattern flow on with null fields rather than failing the query, and everything before `stats` disappears afterwards — only the aggregates and the grouping keys survive.

go deeper

for a junior

Know that plain text needs parse before you can aggregate it, and that stats with a by clause is how you get counts per key. Be able to write one working example.

for a middle

Explain both parse forms, why unmatched lines produce null fields rather than errors, and what bin() adds to a stats grouping to turn a table into a time series.

for a senior

Show the two-query workflow — aggregate to localise, then filter to read raw lines — and say when the answer is to fix the log format upstream instead of maintaining the parse.

for a principal

Frame it as a logging-contract decision: every fragile parse pattern in an incident runbook is a signal that the fleet should be emitting structured logs or a metric instead.

## Why parse exists Logs Insights auto-discovers fields only from JSON log events. A plain-text line like ``` 2026-03-01T10:00:00Z level=ERROR path=/checkout latency=812ms ``` is, as far as the engine is concerned, a single opaque string in `@message`. `parse` is the command that manufactures fields out of that string for the rest of the pipeline. The fields it creates are **ephemeral**: they exist for this query only, nothing is written back to the stored log event. ## The two parse forms **Glob form** — you write the literal text with `*` wherever a value sits, and name the captures in order: ``` parse @message "path=* latency=*ms" as path, latency ``` Each `*` is non-greedy and the literal text between the stars must match exactly. This is readable and is the right default when the log format is stable. **Regex form** — you supply a regular expression with named capture groups: ``` parse @message /path=(?<path>\S+)\s+latency=(?<latency>\d+)ms/ ``` Use this when fields are optional, separators vary, or you need character-class precision. It is slower to write and slower to read, so reach for it only when the glob cannot express the shape. Values captured by `parse` arrive as text, but Logs Insights coerces numeric-looking values when you use them arithmetically or in an aggregate, so `pct(latency, 95)` works on a captured `812` without an explicit cast. ## What happens to lines that do not match This is the single most common surprise. A non-matching event is **not** dropped and does **not** fail the query — it simply flows down the pipeline with the parsed fields null. In an aggregate that shows up as a phantom group whose key is empty, or as a count that is lower than you expected. The fix is an explicit filter, either before the parse (`filter @message like /latency=/`) or after it (`filter ispresent(latency)`), so you are aggregating only over the lines you meant. ## Aggregating with stats `stats` takes one or more aggregate functions and an optional `by` clause: ``` parse @message "path=* latency=*ms" as path, latency | filter ispresent(latency) | stats pct(latency, 95) as p95, avg(latency) as mean, count(*) as n by path, bin(5m) | sort p95 desc ``` The functions you will use most are `count(*)`, `count_distinct(field)`, `sum`, `avg`, `min`, `max`, `stddev` and `pct(field, n)` for percentiles. `count(*)` counts rows; `count(field)` counts rows where that field is present, which is a genuinely different number after a partial parse. ## bin() and the time dimension Without a time key, `stats` gives you one row per group for the whole selected range. Grouping by `bin(5m)` rounds each event's `@timestamp` down to a five-minute boundary and makes that boundary part of the group key, which is what turns a table into a time series — and it is what lets the console render the result as a line chart instead of a table. The bucket size is yours to choose (`bin(1m)`, `bin(1h)`): too fine over a long range produces thousands of rows, too coarse hides the spike you are hunting. ## stats is a wall After a `stats`, the raw events are gone. The only fields that still exist downstream are the aggregates you named with `as` and the keys you grouped by. So `sort` and `limit` after a `stats` operate on aggregate rows — `| sort p95 desc | limit 10` gives you the ten worst paths — but you cannot go back and print `@message` for the offending events in the same query. The normal workflow is two queries: aggregate to find the guilty path or host, then a second query filtered to it to read the actual lines. ## When not to parse at all If you find yourself maintaining a fragile `parse` pattern for a question you ask every week, the real fix is upstream: emit the log as JSON so the fields are discovered automatically, or publish the number as a metric so the recurring question is answered without scanning logs at all. `parse` is an incident tool, not an architecture.

  • Your parse pattern matches most lines but the result contains one group with an empty key. What happened?
    Some events did not match the pattern, so their parsed fields are null, and `stats ... by` treats null as its own group. Nothing failed — parse is lenient by design. Add `filter ispresent(field)` after the parse, or a `filter @message like /.../` before it, so only the lines you intend to measure reach the aggregate.
  • After computing a p95 per path you want to read the slowest actual requests. Can you do that in the same query?
    No. `stats` collapses the event stream, so `@message` no longer exists downstream — only the aggregates and grouping keys. Run a second query filtered to the path you identified, sorted by the parsed latency descending with a small limit. Aggregate first to find where, then drill in to read what.
  • When would you use the regex form of parse instead of the glob form?
    When the shape is not a fixed sequence of literals: optional fields, variable separators, or values that need a character class to delimit cleanly. Named capture groups also let you pull fields out of order. The glob form is more readable and should stay the default for stable formats such as key=value logging.

saying these in an interview costs you the question

  • Expects a non-matching line to fail or skip the query
  • Tries to display @message after a stats command
  • Thinks parse writes fields back to the stored log event
  • Groups by path only and wonders where the time series went
  • Believes parsed values cannot be used in numeric aggregates

context