skip to content

Why does adding a high-cardinality label to a Grafana Loki stream destroy its cost advantage?

level: seniorimportance: must knowfreq 64%

answer

  1. one label value, one more of something
  2. the multiplier is on streams, not columns
  3. tiny chunks flushed on idle rather than size
  4. ingester memory scales with active streams
  5. identifiers belong in a line filter instead

basics

~20 s

Each distinct label value creates a separate Loki stream with its own chunks. A high-cardinality label multiplies streams, so the index swells, ingesters hold thousands of barely-filled chunks in memory, and object storage fills with tiny, poorly-compressed files.

solid answer

~50 s

Loki's index is sized by the number of streams, and a stream is one unique label set — so a label with thousands of values does not add a column, it multiplies the stream count. Each new stream gets its own open chunk in an ingester, which costs memory whether or not anything else is ever written to it, and which is flushed on idle time or age rather than on being full. The result is a bloated index plus a swarm of tiny chunks that compress badly and must each be fetched at query time. Past a per-tenant active-stream limit, writes for new streams are simply rejected. **The fix is to keep identifiers in the line and extract them at query time**: a line filter such as `|= "WO-48213"` or a parser stage like `| logfmt` costs a scan you were already paying for, and costs nothing at ingest.

code

logql · 5 lines
logql
{app="planner", env="prod", work_order_id="WO-48213"}

{app="planner", env="prod"} |= "WO-48213"

{app="planner", env="prod"} | logfmt | work_order_id="WO-48213"

go deeper

for a junior

Recall the rule rather than the mechanics: a label with many possible values is dangerous in Loki, and ids belong inside the log line where a filter can find them.

for a middle

Explain why. Each distinct label value is a separate stream with its own chunk, so cardinality multiplies streams, and chunks that flush on idle instead of on size compress badly.

for a senior

Diagnose it under load. Connect stream churn to ingester memory, to small-object pressure on storage, to rejected pushes at the tenant stream limit, and describe how you would find the offending label and roll it back.

for a principal

Own the guardrails, not the incident. Decide what the tenant limits should be, who may add a label, and how teams get per-request searchability without buying it with ingest capacity everyone shares.

## What a new label value actually creates In Loki a **stream** is one unique set of label key/value pairs, and every stream owns its own chunks. So a label is not a column you attach to rows — it is a partitioning key for physical storage. Adding `level` with four values to a set that produced 492 streams can produce close to four times as many. Adding `work_order_id`, which takes a new value on every job, produces a new stream per job forever. Take a maintenance-planning platform on 41 hosts running 12 containers each. With `{app, env, host, level}` it writes roughly 492 to 2,000 streams — an index measured in megabytes. A well-meaning change adds `work_order_id` to make individual jobs greppable. The platform books about 71,400 work orders a week. The stream count does not grow by one; it grows by 71,400 a week, and each of those streams carries perhaps nine log lines in its whole life. ## The three costs, in the order they hurt 1. **Ingester memory.** Every active stream holds an open chunk in memory on every ingester replica that owns it. Thousands of near-empty chunks cost far more per stored byte than a few full ones, and the ingester's working set is driven by stream count rather than log volume. 2. **Chunk flushing.** A chunk is closed when it reaches its target size, when the stream falls idle, or when it has been open too long — the ingester's `chunk_target_size`, `chunk_idle_period` and `max_chunk_age` settings between them. A stream with nine lines never reaches the target size, so it flushes on idle or age as a tiny object. Compression works on repetition within a chunk, so tiny chunks compress badly, and object storage accumulates millions of small files where it should hold thousands of large ones. 3. **Index and query cost.** The index grows with the stream count, so it stops fitting comfortably in cache. Worse, a query whose selector matches many of those streams must resolve and fetch every one of their chunks, paying a per-object round trip for nine lines each. The design's promise — a small index and a big sequential scan — inverts into a large index and a scattered one. Past all of that sits a hard wall: Loki enforces a per-tenant limit on active streams, and once it is crossed, pushes creating new streams are rejected. The failure mode is not gradual degradation but log loss, usually noticed by whoever is on call rather than by the team that added the label. ## Label or filter? The decision table | Property of the value | Put it in | Why | |---|---|---| | Small, bounded, known set (`env`, `app`, `cluster`, `namespace`, `level`) | A stream label | You will select on it, and it partitions storage usefully | | Unbounded or per-request (work order id, trace id, user id, session id, url path) | The log line | Extract at query time with a line filter or parser; costs nothing at ingest | | Changes on every deploy (image digest, pod name suffix, build number) | The log line, or drop it | Each deploy would otherwise churn the entire stream population | | Numeric measurement (duration, byte count, queue depth) | The log line | Continuous values are the worst case for stream count | The query-time equivalent is genuinely cheap because you are already paying for the scan. `{app="planner", env="prod"} |= "WO-48213"` narrows by the bounded labels, then matches the raw bytes of exactly those chunks. A structured line can go further: `| logfmt` or `| json` turns fields into labels **for the duration of the query only**, so `| work_order_id="WO-48213"` filters without a single extra stream ever existing at rest. ## Living with a mistake already made The damage is retroactive in one direction only. Removing the label from the agent stops new streams appearing, but every chunk already flushed keeps the label set it was written with, so the index stays fat until retention ages it out. Practical steps, roughly in order: - Find the offender: list the streams under the suspect selector and look for a label whose value count tracks request volume rather than fleet size. - Drop the label at the collection agent, not in the query — the cost is at ingest. - Check the discovery meta labels the agent attaches automatically; those are a common accidental source, and dropping them at relabel time is a one-line change. - Expect a period where old and new label sets coexist in queries, and write selectors that tolerate both until retention clears the old ones. The rule to state in an interview is short: **labels answer "which logs", lines answer "which one".** Anything you would only ever search for once belongs in the line.

  • A team wants per-customer dashboards and proposes a customer_id label. How do you answer?
    Ask how many customers there are and whether the number is bounded. A few dozen fixed tenants is a legitimate label and partitions storage the way you want. An open-ended and growing customer base is stream churn wearing a business justification, so keep the id in the line and build the dashboards on a parser stage, or route genuinely separate tenants through Loki's multi-tenancy instead of encoding them as a label on one shared tenant.
  • You removed the offending label an hour ago and the index is still huge. Why?
    Chunks are immutable once flushed and carry the label set they were written with, so the streams already in object storage do not disappear when the agent configuration changes. New writes land in the smaller stream population immediately, while the old population ages out only under retention. Until then queries can match both shapes, so selectors have to tolerate the old label being present on historical data and absent on recent data.
  • Does a high-cardinality label hurt more at write time or at read time?
    Write time is where it becomes an outage. Ingester memory and the per-tenant active-stream limit are hard resources, and crossing the limit rejects pushes and loses logs outright. Read time degrades badly but gracefully: queries fetch many small objects instead of few large ones and get slower. Interviewers look for that ordering, because the instinct is to describe it as a query performance problem when the first thing it actually breaks is ingest.

saying these in an interview costs you the question

  • Treats a Loki label and a line filter as interchangeable
  • Puts request, trace or work-order ids in labels to make them searchable
  • Thinks stream cardinality only costs disk, not ingester memory
  • Believes removing the label immediately shrinks the existing index
  • Says compression makes the number of streams irrelevant
  • Assumes Loki degrades gradually rather than rejecting pushes