skip to content

In a log store that indexes only stream labels and scans the lines, which query shapes are fast and which are catastrophic?

level: middleimportance: should knowfreq 52%

answer

  1. Only the identifying key/value pairs are indexed
  2. Select streams, select blocks, then scan
  3. Cost is bytes decompressed, not matches
  4. Filter complexity is not the cost driver
  5. Unbounded selection over a long window kills

basics

~20 s

Only the labels identifying each stream are indexed, so a query selects streams and a time range, then fetches, decompresses and scans those lines. Narrow labels over a short window are fast; a long window with no label matcher is catastrophic.

solid answer

~40 s

The index holds only a small set of key/value labels - service, environment, host - identifying each stream of lines, plus which compressed blocks cover which time ranges. Everything inside a line is unindexed. A query therefore runs in three stages: **match labels** to select streams, **select blocks** overlapping the time window, then **fetch, decompress and scan** them line by line. Cost is proportional to the bytes stages one and two failed to eliminate - not to how many lines match and not to how complex the filter is. So an elaborate expression over one service for thirty minutes is cheap, while a trivial substring match with no useful label matcher over a week has to read essentially everything. The right optimisation is always narrowing the selection, never simplifying the filter.

code

pseudocode · 7 lines
pseudocode
streams = label_index.match(service="booking-api", env="prod")
blocks  = block_index.overlapping(t0, t1, streams)

for block in blocks:                    # fanned out across workers
    for line in decompress(block):      # cost lives here
        if "ferry_id=48213" in line:
            emit(line)

go deeper

for a junior

Be ready to say that this kind of store indexes only the few labels identifying where lines came from, and that everything else is found by reading the lines themselves at query time.

for a middle

Explain the three query stages and why cost tracks the compressed bytes fetched. Give a concrete fast shape and a concrete catastrophic shape rather than saying it depends.

for a senior

Show you have run one under pressure: per-query limits on how much data may be touched, protecting the write path from concurrent scans, and refusing the label-promotion workaround engineers ask for.

for a principal

Own the estate-level consequence: this design assumes searches begin from a known service and window, so it must be paired with conventions and correlation identifiers that make that assumption true.

## What such a store indexes, and what it deliberately does not In this family the store keeps a small **index of log-stream labels** - a handful of key/value pairs such as service, environment, host or namespace that identify a stream of lines - together with enough bookkeeping to know which compressed blocks of lines belong to which stream over which time range. The lines themselves are stored close to how they arrived, compressed in blocks, and nothing *inside* a line is indexed. There is no structure anywhere that can answer *which lines contain this token*. A query therefore runs in three stages: 1. **Select streams.** Evaluate the query's label matchers against the label index. This produces a set of streams and nothing else. 2. **Select blocks.** For those streams, find the compressed blocks whose time ranges overlap the requested window. 3. **Scan.** Fetch those blocks, decompress them, and apply the rest of the query - text filters, field extraction, counting - line by line, usually fanned out across many workers. Everything about performance follows from stage three: **query cost is proportional to the compressed bytes that stages one and two failed to eliminate.** Not to how many lines match. Not to how complex the filter is. To bytes fetched and decompressed. ## The fast shape and the catastrophic shape | Query | What it touches | Verdict | | --- | --- | --- | | Narrow label matchers, 30-minute window, text filter | One stream's blocks for half an hour | Fast, and the filter is nearly free | | Narrow label matchers, 7-day window | One stream's blocks for a week | Slow but bounded; parallelism helps | | Broad label match, 30-minute window | Many streams, short range | Survivable; costs cluster CPU | | No effective label matcher, 7-day window, text filter | Every block of every stream in the window | Catastrophic | On a ferry-timetable booking platform ingesting 1.7 TB of logs a day, a filter for a booking reference inside one service over 30 minutes might decompress 4.2 GB. The identical filter with no service matcher over seven days has to work through something near 11.9 TB - and 71% of that belongs to one team's schedule-sync service, which is almost certainly not the service you were investigating. Same filter, three orders of magnitude apart in cost, and the expensive version competes with every other query and with the write path for the same CPU and read bandwidth. The corollary that catches people: **complexity is not the cost driver.** An elaborate regular expression over a narrow selection is cheap; a trivial substring match over an unbounded selection is ruinous. Optimising the filter is the wrong instinct - narrowing the selection is the right one. ## Why anyone chooses this - **Writes are cheap and content-independent.** There is no analysis step, so write capacity scales with bytes, not with what is in them. A service that starts logging a new format costs nothing extra. - **Stored bytes are close to compressed raw bytes.** There is no derived structure multiplying the footprint, which is what makes keeping everything affordable. - **You pay only for the questions actually asked.** In an estate where most searches begin with *which service, what time window*, most of the stored data is never read at all. - **Onboarding is trivial.** No schema to agree, no mapping to maintain, no coordination between teams about field names before their logs are useful. ## The traps - **The label temptation.** When a query is slow, the tempting fix is to promote the thing you filter on into a label. A per-request or per-customer value creates one stream per distinct value, and the number of streams is exactly the axis this design is sensitive to: it inflates the index it was trying to keep small and fragments the data into many tiny, poorly compressing blocks. The discipline is that labels describe **where logs came from**, never **what happened inside them**. - **Concurrency is a capacity problem, not a latency problem.** Ten engineers each running a week-wide scan during an incident is a fleet-wide CPU and read-bandwidth event. Limits on how much data a single query may touch are normal operational hygiene, not an insult to users. - **Parallelism is not a saving.** Spreading a scan across more workers reduces wall-clock time; it does not reduce bytes read, so the bill and the contention are unchanged. - **Retrospective questions are the weak spot.** *Has this error string ever appeared anywhere?* is precisely the query the model is worst at, and it is a question incidents produce constantly. Said in one sentence: this design moves cost from write time to read time, so ingesting everything is cheap and asking a question that cannot be narrowed by labels and time is expensive.

  • A search is slow, so an engineer wants to make the value they filter on a stream label. What do you tell them?
    That it fixes their query and damages the store. Labels define streams, so a per-request or per-customer value creates one stream per distinct value; the small index this design depends on inflates, and the data fragments into many tiny, badly compressing blocks. Labels should describe where logs came from - service, environment, host - never what happened inside them. The right fix is a narrower time window or a more specific existing label.
  • Does running the scan across more workers reduce its cost?
    It reduces wall-clock time, not cost. The same compressed bytes are still fetched and decompressed; the work is just spread wider, so the read bandwidth and CPU consumed are unchanged and the contention with other queries and with ingest is, if anything, more intense. Limiting how much data one query may touch is the control that actually protects the platform.
  • Why is a complex regular expression cheap here but a broad time range expensive?
    Because the byte volume decides the cost. Once blocks are being decompressed, evaluating a more elaborate expression on each line is marginal work. Widening the window or dropping a label matcher, by contrast, multiplies the number of blocks that must be fetched and decompressed - which is the dominant term. Optimise the selection, not the expression.

It is a warehouse where only the shelf labels are catalogued: fetching one labelled crate and rummaging through it is quick, but finding an item with no idea which shelf means opening every crate in the building.

saying these in an interview costs you the question

  • Thinks the store can look up a token without scanning
  • Believes a simpler filter makes an unbounded query affordable
  • Promotes request identifiers into labels to speed searches up
  • Assumes more query workers reduce the bytes read
  • Cannot say what the label index actually contains