skip to content

In Loki's LogQL, why can only the stream selector make a query cheap?

level: middleimportance: should knowfreq 58%

answer

  1. a query has two distinct halves
  2. one half talks to the index
  3. the rest runs on lines already fetched
  4. line filters are cheaper than parsers
  5. narrow the selector and the time range first

basics

~20 s

The stream selector decides which chunks are read from object storage. Every filter, parser and metric stage after it runs over lines already fetched and decompressed, so they change what you see, not how much is read.

solid answer

~50 s

A LogQL query has two halves. The **stream selector** — `{app="planner", env="prod"}` — plus the time range resolves, through the index, to a set of chunks; that is the only decision that changes how many bytes leave storage. Everything after it is a **pipeline** running over the lines already retrieved: line filters (`|=`, `!=`, `|~`, `!~`), parsers (`| logfmt`, `| json`, `| pattern`, `| regexp`), label filters on the fields a parser extracted, and formatters. Wrapping the whole thing in `rate(...)` or `count_over_time(...)` turns it into a metric query, but the input volume is still whatever the selector chose. Loki requires the selector to contain at least one matcher that does not match the empty string, precisely so a query cannot ask for everything. Inside the pipeline, order still matters for CPU: put cheap line filters before parsers, since parsing every line costs far more than matching bytes.

code

logql · 3 lines
logql
sum by (host) (
  rate({app="planner", env="prod"} |= "dispatch_failed" | logfmt | attempts > 3 [5m])
)

go deeper

for a junior

Recognise the shape of a LogQL query: a label selector in braces, then optional filters after a pipe. Know that the braces part is mandatory and cannot be left empty.

for a middle

Explain the split. The selector plus time range resolves through the index to chunks; filters, parsers and label filters run over lines already fetched, and line filters are far cheaper than parsers.

for a senior

Tune a real query in front of the interviewer: narrow the selector, shorten the range, order line filters before parsers, and say why a heavily-filtered query can still be the one overloading the read path.

for a principal

Set the conventions that make good queries possible: which labels every service must carry, whether lines are structured enough to parse, and what default time ranges dashboards are allowed to use.

## The two halves Read any LogQL query as `{selector} <pipeline>`, optionally wrapped in a metric function. ``` sum by (host) ( rate({app="planner", env="prod"} |= "dispatch_failed" | logfmt | attempts > 3 [5m]) ) ``` - `{app="planner", env="prod"}` is the **stream selector**. Combined with the query's time range it goes to the index, resolves to a set of streams, and from those to the chunks that must be downloaded and decompressed. - `|= "dispatch_failed" | logfmt | attempts > 3` is the **pipeline**. It runs line by line over whatever came back. - `rate(... [5m])` and `sum by (host)` turn the surviving lines into a numeric series. Only the first part touches the index. This is the single most important operational fact about LogQL, and the reason a query that "already filters heavily" can still be the one melting the read path. ## What each stage costs | Stage | Example | Runs on | Relative cost | |---|---|---|---| | Stream selector | `{app="planner", env="prod"}` | The index | Decides bytes fetched | | Line filter | `\|= "dispatch_failed"` | Raw bytes of each line | Cheapest per line | | Parser | `\| logfmt`, `\| json` | Every line that survived | Expensive per line | | Label filter | `\| attempts > 3` | Fields a parser extracted | Cheap, but requires the parser | | Formatter | `\| line_format`, `\| label_format` | Surviving lines | Cheap, cosmetic or reshaping | | Metric function | `rate(...)`, `count_over_time(...)` | Surviving lines | Aggregation over the result | Two consequences follow. First, **adding a filter never makes a query cheaper than narrowing the selector**, because the filter runs after the bytes have already been paid for. Second, **within the pipeline, order changes CPU**: a line filter that discards 98% of lines before `| logfmt` means the parser runs on 2% of the data. Writing `| logfmt |= "dispatch_failed"` instead parses everything and then throws it away. ## Why the selector is mandatory and why it must be selective Loki requires the stream selector to include at least one matcher that cannot match the empty string, so a query cannot ask for the entire estate by omission. That protects against the accident, not against the choice: `{env="prod"}` satisfies the rule and can still mean every chunk from every application in the environment. Time range is the selector's silent partner. The read path splits a query by time interval and fans the sub-queries out across queriers, so a well-scoped query over an hour is genuinely parallel work on a small set of chunks. The same query over thirty days multiplies the chunk set by roughly thirty. When someone reports that Loki is slow, the first two things to look at are the selector and the range, in that order — the pipeline is almost never the cause. ## Working the query, in practice For a maintenance-planning platform on 41 hosts running 12 containers each, an on-call engineer with a failing dispatch typically narrows in this order: 1. **Select as tightly as the labels allow.** `{app="planner", env="prod", host="turbine-gw-03"}` if the host is known — every label added here removes chunks from the scan. 2. **Shorten the range.** Fifteen minutes around the incident, not the default hours-long window. 3. **Line filter next.** `|= "WO-48213"` matches raw bytes and is the cheapest way to cut the surviving set. 4. **Parse only what is left.** `| logfmt` or `| json` on the remainder, then label filters such as `| attempts > 3` on the extracted fields. 5. **Aggregate last, if at all.** `count_over_time` or `rate` over the pipeline output when the question is "how often", not "which line". A note on the labels a parser produces: they exist only for the life of the query. `| logfmt | work_order_id="WO-48213"` is filtering, not indexing, and nothing about it creates a stream. That is exactly why identifiers belong in the line rather than in the selector — the query language gives you the searchability without the storage cost. ## The sentence to say in an interview **The selector decides how much data the query reads; everything after it decides how much of that data you look at.** A candidate who can say that, and can then explain that line filters should precede parsers because parsing is the expensive per-line operation, has demonstrated they have actually tuned a slow query rather than only written one.

  • Does moving a line filter before a parser in a LogQL pipeline change the result?
    Not if the filter matches the raw line rather than a parsed field — it changes only the cost. A line filter matches bytes and needs no structure, so running it first means the parser only sees the lines that survived. Order does change the result when the filter depends on something the parser produced: a label filter on an extracted field has to come after the parser that extracted it, and putting it earlier is a query error rather than a slow query.
  • A LogQL query is slow and the pipeline is already a single cheap line filter. Where do you look?
    At the selector and the time range, because those are the only things that determine how many chunks are fetched. Count how many streams the selector matches and how many hours the range covers; a broad selector over a long window is scanning far more than the filter suggests. Look for a bounded label that could be added, and narrow the range around the event. Only after that is it worth asking whether the read path has enough queriers.
  • Labels created by a LogQL parser stage look like stream labels. Are they?
    No, and the distinction matters. A parser such as `| logfmt` or `| json` extracts fields into labels that exist only while the query runs, so you can filter and aggregate on them freely without any cost at ingest. Stream labels, by contrast, are chosen by the collection agent, are written into the index, and partition storage permanently. Parser-produced labels are the reason high-cardinality values can safely live in the log line.

saying these in an interview costs you the question

  • Thinks adding a line filter reduces the data fetched from storage
  • Cannot name the two halves of a LogQL query
  • Puts a parser before the line filter and calls it equivalent
  • Believes parser-extracted labels are written into the index
  • Ignores the time range when explaining why a query is slow