skip to content

How do you choose between a full-text-indexed log store and a label-indexed scanning one, and what later says you chose wrong?

level: seniorimportance: should knowfreq 48%

answer

  1. Read the query log before the product page
  2. Known service and window, or open-ended
  3. Volume skew decides what indexing wastes
  4. Watch for shortened windows and shadow copies
  5. Route by log class, not one platform decision

basics

~20 s

Decide from the query log, not the product page: if almost every search already knows the service and time window, buy cheap writes and pay per query; if searches are open-ended or aggregate across everything, buy the index.

solid answer

~40 s

Answer four questions first. **What do people actually search?** Sort a month of real queries into needle-in-a-known-haystack versus open-ended discovery and aggregation. **Who queries, and how automatically?** Alerting on log content and support staff searching identifiers force an index; on-call engineers arriving with a service and a timestamp do not. **What is the volume and its skew?** If one team is most of your bytes and nobody searches them, indexing them is pure loss. **How structured are the logs, and can you change the producers?** Then watch for the reversal signals: write-path lag with unqueried indexed fields, or shortened windows, abandoned searches and engineers promoting identifiers into labels. The strongest signal either way is teams keeping a second copy of the logs somewhere else to analyse them.

go deeper

for a junior

Be ready to state the trade in one line - pay while writing for fast search later, or write cheaply and pay while searching - and to say that how people search should decide it.

for a middle

Explain which workloads suit each model and name concrete inputs: query patterns, concurrency, daily volume and its skew, and how structured the logs already are.

for a senior

Bring evidence and operations: audit real queries, replay a real day including the dominant producer, and name the symptoms that say the choice has gone wrong before the bill does.

for a principal

Own the strategy: route by log class rather than choosing once, keep the raw records so the decision stays reversible, and pick in advance which failure mode your on-call rotation can survive.

## Answer these before you look at any product The choice is not between two products, it is between paying at write time and paying at read time. Four pieces of evidence decide it, and only one of them is a benchmark. 1. **What do people actually search for?** Not what they say - what the query logs show. Sort last month's searches into two piles. Pile A: *this service, this window, find the request*. Pile B: *has this string ever appeared anywhere*, *count these by field across the estate*, *show me the distribution*. Pile A is a needle in a **known** haystack and a scan-based store serves it well. Pile B needs an index built in advance. 2. **Who queries, how often, and how concurrently?** Five on-call engineers who arrive with a service name and a timestamp behave completely differently from thirty support agents searching order references, or from an alerting system evaluating text patterns across everything every minute. Automated readers are the ones that quietly make an index mandatory, because they run whether or not anyone is watching. 3. **What is the volume and its shape?** Total bytes per day; how far back queries realistically reach; and above all the skew. On a ferry-timetable booking platform ingesting 1.7 TB a day, one team's schedule-sync service is 71% of the volume - and if nobody searches that 71%, indexing it is the largest single line item on the bill for no return. 4. **How structured are the logs already, and who can change them?** Consistent structured records across services make the write-time model cheap to adopt. A mix of vendor components, legacy services and free text means the up-front declaration will be wrong, and being wrong is permanent for data already written. The rule of thumb that falls out: **if nearly every query already knows the service and the window, buy cheap writes and pay per query. If queries are open-ended discovery or aggregation over everything, buy the index.** ## The signals that say you chose wrong | Symptom | What it means | | --- | --- | | Write path lags at every traffic peak while CPU saturates | You are indexing content nobody searches | | Query audit shows a handful of fields carry almost all searches | Most of the structure you pay for is dead weight | | Retention is being cut to afford the platform | The write-time bill, not the data, is the constraint | | Engineers shorten windows until a search finishes | Read-time cost has outgrown the workload | | Searches are abandoned or time out during incidents | The store cannot answer the questions incidents ask | | Teams keep a second copy of logs elsewhere to analyse | The strongest signal of all; users have routed around you | | People are promoting high-cardinality values into labels | The scan model is being bent into an index, badly | The last one deserves emphasis. When engineers start adding per-request or per-customer identifiers as stream labels to make their searches fast, they are telling you the read-time model does not fit their questions - and the workaround damages the write path, because stream count is the axis a label-indexed design is sensitive to. ## What a senior answer adds - **Do not choose once for all logs.** The realistic architecture routes by class: a small, high-value, fully searchable set - request-level events, errors, audit and security-relevant records - and a large, cheap, scan-only bulk for debug chatter and the dominant machine-generated stream. That is one decision per log class, not one decision per platform. - **Instrument the decision.** Sample and retain query logs from day one. Without them, every future argument about the platform is anecdote, and the audit in signal two above is impossible. - **Migration is not symmetric.** Moving to a write-time model gives you nothing retroactively - the old data stays unindexed - so you run both for the retention window. Moving away is easier, because the raw records were kept. - **Pilot on the real skew.** Benchmarks on synthetic uniform logs mislead badly when one producer is most of your volume. Replay a real day, including the 71% stream, before committing. - **Name the failure you can live with.** The write-time model fails by delaying fresh logs during incidents; the read-time model fails by making broad searches slow or refused during incidents. Both hurt at three in the morning. Decide in advance which one your on-call rotation can work around, because you will meet it.

  • You have no query logs and must decide this quarter. What do you do?
    Start collecting them immediately, and in the meantime interview the two populations that matter: whoever is on call, and whoever answers customer questions. Ask each to walk through their last five real searches. Then pilot on a replay of a real day rather than synthetic uniform logs, because a dominant producer distorts every benchmark. Decide for one log class, not the whole estate, and keep the raw records so the decision stays reversible.
  • Why is migrating toward a write-time indexed store harder than migrating away from one?
    Because indexing is not retroactive. The new store can only index what arrives after the cutover, so old questions still have to be asked of the old system and you run both for the length of the retention window. Migrating away is easier: the raw records were kept, so they can be shipped into the scan-based store and remain answerable, just more slowly.
  • One team produces most of your log volume and never searches it. What is the move?
    Treat it as its own class. Keep storing it - it is cheap and it is evidence - but stop paying to make it searchable: exclude it from indexing, or route it to the scan-based path while the high-value classes stay indexed. Then push back on the producer about log level and duplication, because the cheapest byte is the one never shipped.

saying these in an interview costs you the question

  • Picks a product before looking at real query patterns
  • Benchmarks on synthetic logs with no volume skew
  • Assumes one storage model must serve every log class
  • Ignores automated readers such as content-based alerting
  • Treats the decision as reversible without keeping raw records
  • Cites ingest price alone while query cost is unmeasured