skip to content

In Grafana Loki, what is a stream, what goes into the index, and what does not?

level: juniorimportance: must knowfreq 76%

answer

  1. only part of a log entry is indexed
  2. labels in, line text out
  3. a unique label set defines one stream
  4. lines become compressed chunks in object storage
  5. index maps label sets to chunk references

basics

~20 s

Loki indexes only the label set attached to each log line, never the line's text. A stream is one unique combination of label values; its lines are batched into compressed chunks kept in object storage.

solid answer

~50 s

A **stream** in Loki is one unique set of label key/value pairs: `{app="planner", host="turbine-gw-03", level="error"}` is a different stream from the same set with `level="warn"`. Loki's index maps those label sets to the chunks holding their entries over a time range, and it holds no terms from the log lines themselves. Incoming entries for a stream are appended to a chunk held in memory by an ingester, then compressed and flushed to object storage once the chunk is big enough, has gone idle, or has simply been open too long. A query therefore has two separate costs: a small index lookup, and a scan that downloads and decompresses whatever chunks the selector matched. Writes stay cheap because nothing is tokenised at ingest; reads stay affordable only while the label selector keeps the scanned volume small.

code

json · 11 lines
json
{
  "streams": [
    {
      "stream": { "app": "planner", "env": "prod", "host": "turbine-gw-03" },
      "values": [
        [ "1757068800000000000", "level=error msg=\"dispatch failed\" work_order=WO-48213" ],
        [ "1757068800417000000", "level=info msg=\"retry scheduled\" work_order=WO-48213" ]
      ]
    }
  ]
}

go deeper

for a junior

Be ready to state the one-line rule: labels are indexed, log text is not, and a unique label set is a stream. Know that lines end up compressed in object storage rather than in a searchable database.

for a middle

Explain the mechanics: how an ingester accumulates a per-stream chunk and what triggers a flush, what the index actually maps to what, and why a query has a lookup cost and a separate scan cost.

for a senior

Show you have operated it. Talk about how label choices made at collection time are frozen into flushed chunks, how stream count sizes the index, and how you keep scans bounded when a team asks for a fleet-wide search.

for a principal

Own the trade itself: Loki buys cheap ingest by deferring work to query time, and the labels are the only lever anyone has over that cost. Decide when that trade suits your estate and when a full-text store is worth its ingest bill.

## The bargain Loki makes Loki stores two very different things in two very different ways. The **labels** attached to a log entry go into a small index. The **line itself** goes into a compressed chunk in object storage, and is never tokenised, analysed or indexed at write time. Almost everything interesting about operating Loki — why ingest is cheap, why a badly-scoped query is slow, why one careless label ruins both — follows from that one split. Contrast it with the alternative in one sentence, because interviewers usually ask for it: a full-text log store analyses every line at ingest and builds term-level structures, so any word is a fast lookup but ingest costs CPU and the index rivals the data in size. Loki refuses to do that work, and pays for the refusal at query time instead. ## What a stream is A **stream** is one unique set of label key/value pairs. `{app="planner", host="turbine-gw-03", level="error"}` and `{app="planner", host="turbine-gw-03", level="warn"}` are two streams, because one value differs. Three things follow: - A stream is **not** a file, a container or a process. It is whatever label set the sending agent chose to attach, and two containers with identical labels write into the same stream. - Every entry belongs to exactly one stream and carries a nanosecond timestamp plus the raw line. There are no other per-entry fields at rest — structure inside the line is the line's business. - The number of distinct streams, not the number of log lines, is the quantity that sizes the index. ## What the index holds, and what it does not | Artefact | Contains | Where it lives | Grows with | |---|---|---|---| | Index | Label sets, and references to the chunks that hold each stream's entries over a time range | Object storage, downloaded and cached by the read path | Number of distinct streams | | Chunk | Compressed, timestamp-ordered entries for exactly one stream | Object storage | Volume of log text | No word from a log line appears in the index. Searching for a work-order id is therefore not a lookup at all — it is a scan of every chunk the selector matched, decompressed and pattern-matched line by line. Loki makes that fast by brute force and parallelism: the read path splits the query by time, fans the chunks across many queriers, and filters compressed blocks in parallel. ## The write path 1. An agent tails a source, decides the label set, and pushes batched entries to Loki's push endpoint. 2. A distributor validates the batch, hashes the label set to choose ingesters, and replicates the write. 3. An ingester appends each entry to that stream's open in-memory chunk. 4. When the chunk hits its target size, the stream falls idle, or the chunk has been open too long, it is compressed and flushed to object storage and the index gains a reference to it. Because the chunk is per-stream, the labels chosen in step 1 decide the physical layout of everything written in step 4. That is why label choice is an ingest-time decision with permanent consequences: the data already flushed keeps the streams it was written with. ## The read path A query names a time range and a **stream selector** such as `{app="planner", env="prod"}`. Loki consults the index for the streams matching that selector inside the range, resolves them to chunk references, fetches and decompresses those chunks, and only then applies any filtering the query asked for. The selector is the only part of the query that changes how many bytes come off object storage; everything after it changes only what survives. For a small estate — say a maintenance-planning platform on 41 hosts running 12 containers each — a label set of application, environment, host and severity yields a few hundred streams. The index for that is trivial and the chunks compress well, because each stream produces a steady flow of similar lines that fill a chunk before it is flushed. ## Where the cheapness ends The design pays off only when two things hold. First, the stream count stays bounded, so the index stays small and chunks fill up before they are flushed. Second, queries carry a selective label selector, so a scan touches megabytes rather than terabytes. Break either and the model inverts: many rarely-written streams produce a bloated index and a swarm of tiny, poorly-compressed chunks, and a selector like `{env="prod"}` over a wide time range asks the read path to decompress the whole estate to find one line. That is the whole trade in one sentence: **Loki moves work from ingest to query, and gives you the labels as the only lever over how much query work there is.**

  • If no word of the log text is indexed, how does searching for a substring ever finish quickly?
    It does not use an index at all — it is a parallel scan. The stream selector and the time range together decide which chunks are fetched; the read path splits the range into sub-queries, spreads the chunks across many queriers, and matches the substring against decompressed blocks concurrently. Speed comes from having narrowed the byte count first and then throwing parallelism at what is left, which is why an unselective selector cannot be rescued by a clever filter.
  • What happens to a stream's open chunk when the process writing it stops logging?
    It is flushed anyway. An ingester closes a chunk when it reaches its target size, when the stream has been idle for a configured period, or when the chunk has been open past a maximum age — whichever comes first. A stream that produced twelve lines and went quiet still becomes a chunk in object storage, just a tiny and poorly-compressed one. That is why a large population of rarely-written streams is expensive even when the total log volume is small.
  • Where does the index itself live, given that the chunks are in object storage?
    Alongside them. Loki builds index files locally and ships them to the same object-storage bucket as the chunks, and the read path downloads and caches the pieces it needs, often through a dedicated index-gateway component. That is what makes a Loki cluster largely stateless: the durable state is a bucket of chunks plus a bucket of index files, and the compute tiers can be scaled or replaced without moving data.

It is a library that indexes the shelf each book sits on but not a single word inside any book: finding the shelf is instant, and after that you read.

saying these in an interview costs you the question

  • Claims Loki builds a full-text index over log line content
  • Thinks a stream means one file, one container or one process
  • Says log lines are stored uncompressed in a database table
  • Believes more labels always makes queries faster
  • Cannot say where log bodies live after an ingester flushes them