What does a log backend that full-text indexes every line at ingest buy you, and what does it cost?
answer
- The work happens before data is queryable
- Every token becomes searchable later
- Paid on write CPU and stored bytes
- Query cost tracks matches, not total volume
basics
~20 sA full-text index built at ingest makes any word in any line searchable without declaring it first, and makes aggregations fast. You pay twice: CPU on the write path for every record, and index bytes stored alongside the logs.
solid answer
~40 sThis kind of store analyses each record as it arrives, splitting the message into tokens and recording which records contain each one, and keeps declared fields in a typed form. That buys **search for terms nobody anticipated**, aggregations over fields, and response times that track the number of matches rather than the size of the corpus - a term in 300 records costs about the same whether you hold a week or a year. The price is paid on every ingested byte: analysis CPU and memory continuously at the full log rate, plus derived structures on disk that can be a large fraction of, or more than, the raw volume. It also front-loads decisions - what is indexed is fixed before the data lands, and adding a field later only helps records written afterwards.
code
pseudocode · 7 linesfor record in incoming_batch:
for token in analyse(record.message):
term_index[token].add(record.id)
for field, value in record.fields:
field_index[field][value].add(record.id)
typed_store[field].append(value)
document_store.append(record)go deeper
Be ready to say plainly that the store does work up front so that searching later is fast, and that the two prices are processing time while writing and extra disk for the index itself.
Explain the mechanics: analysis into tokens, a lookup from term to records, typed fields for aggregation, and why query time then tracks the number of matches instead of the amount of data held.
Show you have operated one. Talk about write-path lag during volume spikes, auditing which indexed fields are actually queried, and cutting indexing on noisy classes while still storing them.
Own the trade at estate scale: what the platform makes searchable is a budget decision about which future questions you are pre-paying for, and it should be set per log class rather than once for everything.
## What a full-text log store builds while writing A store of this family treats every arriving log record as a document. An **analysis** step splits the message into tokens - words, identifiers, numbers, fragments separated by punctuation - and for each token the store records which records contain it. Fields that arrive already structured (`service`, `level`, `status`, `duration_ms`) are additionally kept in a typed form so they can be compared, ranged over and aggregated. What ends up on disk is therefore two things: the records themselves, and a set of lookup structures whose only job is to answer *which records contain X* without reading the records. The decisive property is **when** that work happens: before the data is queryable. A record is not searchable until it has been analysed, so the write path is doing per-record CPU work at the full rate of your log volume, continuously, forever. ## What the index buys at query time - **Search without prior knowledge.** You can look for a token nobody anticipated - a booking reference, a hostname, half a stack-trace class name - across everything, without having declared it as a field beforehand. - **Response time that tracks matches, not corpus size.** A term appearing in 300 records costs roughly what fetching 300 records costs, whether the store holds a week or a year of data. That property is what makes open-ended investigation feel interactive. - **Aggregation over typed fields.** *Count by status per minute*, *the twenty slowest endpoints*, *how many distinct customers saw the error* are answerable inside the store rather than by exporting data somewhere else. - **Content-based alerting.** A rule that fires when a text pattern appears anywhere in the estate is affordable, because evaluating it is a lookup rather than a scan. ## What it costs | Cost | Where it lands | What it scales with | | --- | --- | --- | | Analysis CPU and memory | The write path, continuously | Bytes and records per second | | Index bytes on disk | Storage, for as long as the data lives | What you chose to index, and how diverse those values are | | Ingest lag under load | Freshness of logs during incidents | Peak burst versus provisioned write capacity | | Up-front decisions | Schema and index configuration | How well you predicted future questions | Two of those surprise teams. **Storage amplification.** You are keeping the original record *and* structures derived from it. How much that adds is entirely a function of what you index: indexing free text plus every field can push total stored bytes to a large multiple of the raw volume, while indexing a handful of fields and leaving the rest merely stored costs very little. It is a dial, not a constant - but somebody has to set it deliberately, and the unconsidered default is usually *index everything*. **Ingest lag.** Analysis capacity is finite. When a bad deploy produces ten times the usual log volume, the write path queues and logs arrive minutes late - precisely during the incident that produced them. Note the asymmetry with a scan-based store: that family degrades on the read path under pressure, this one degrades on the write path, and a degraded write path is much harder to work around while you are on the call. Take a ferry-timetable booking platform pushing 1.7 TB of logs a day into a 7-node cluster, where one team's schedule-sync service produces 71% of the volume. Every one of those bytes is analysed on arrival whether or not anyone ever searches it, and an audit of the last month's searches will typically show a small minority of indexed fields carrying nearly all of the queries. The gap between what you pay to make searchable and what anyone actually searches is the characteristic waste of this model. ## Where the decisions bite 1. **What is indexed is decided before the data lands.** Adding a field later makes *new* records searchable by it; records already written keep the structure they were written with. The practical consequence is that the question you did not anticipate cannot be asked of last week's data. 2. **Near-unique values are the expensive kind.** A field holding a request identifier adds a distinct index entry for essentially every record, and buys you a lookup you could usually have done with a service plus a fifteen-minute window. 3. **Free text is the biggest single lever.** Full stack traces and chatty third-party components are where most of the analysis CPU and most of the index bytes go. Excluding a class of noisy, never-searched logs from indexing while still storing them is the standard first optimisation. 4. **A dominant producer dominates the economics.** When one team is most of the volume, changing what *that* stream indexes moves the platform's cost more than any amount of cluster resizing. The honest summary: this model buys query power in advance, on every byte, whether or not the query is ever asked. That is an excellent trade when many people run open-ended searches and aggregations across the data, and a poor one when almost every search already knows which service and which few minutes it cares about.
- Why does excluding a noisy log class from indexing save more than compressing the stored records harder?Compression only touches stored bytes. Excluding a class from indexing removes the analysis CPU for those records on the write path, the index structures they would have produced, and the memory those structures need - while the records are still stored and still retrievable by service and time. It attacks the cost that scales with content rather than the cost that scales with volume.
- You add a new indexed field today. What can you now ask about last week's logs?Nothing new. Records already written keep the structure they were written with, so the new field is searchable only for data ingested after the change. Answering the question for last week means reprocessing or re-ingesting the original records, if you still have them. That permanence is the main reason teams keep the raw record even after extracting fields from it.
- The write path falls behind during a traffic spike and CPU is pinned. What is the usual cause in this model?Analysis of incoming records. Every byte must be tokenised and folded into index structures before it becomes searchable, so a ten-fold burst in log volume is a ten-fold burst in that work. Queues grow and logs surface minutes late - during the incident that generated them. The levers are reducing what gets indexed, shedding low-value records, or adding write capacity.
It is a book with a full concordance in the back: any word is instantly findable, but somebody had to read the whole book to build it and the concordance is now part of the book's weight.
saying these in an interview costs you the question
- Calls the index a compressed copy of the logs
- Assumes indexing is free because disk is cheap
- Thinks adding an indexed field makes old data searchable by it
- Believes search speed depends on total stored volume, not matches
- Cannot name anything that must be decided before ingest