In a log aggregation pipeline, what does each stage — collect, parse, enrich, route, store — do, and which decisions can no later stage undo?
answer
- Follow one line end to end
- Five stages, one direction of travel
- Collect, parse, enrich, route, store
- Only one stage loses data forever
- Raw text makes reprocessing possible
basics
~20 sCollect reads bytes out of the producing process, parse turns text into named fields, enrich attaches context the emitter lacked, route chooses destinations, store indexes and serves. Collect losses are unrecoverable: an unread line exists in no other copy.
solid answer
~50 sFollow one line. **Collect** gets bytes out of the process — usually a node-level log agent tailing the file a container runtime writes, following rotation and re-joining stack traces and runtime-split fragments into one event. **Parse** turns text into fields: timestamp, level, message, and whatever a pattern extracts. **Enrich** adds context the emitter never had — node, namespace, workload, region. **Route** decides which destinations get which records, including fan-out to several and dropping the rest. **Store** indexes and retains what arrives. The direction is what matters. Collect losses are unrecoverable, because a line the agent never read is gone once the file is rotated or garbage-collected. Parse and enrich mistakes are recoverable *only if* the raw text and the enrichment inputs still exist — which is why keeping the original line beside the parsed fields usually pays for its storage.
code
json · 12 lines{
"timestamp": "2026-03-11T09:14:27.481Z",
"level": "INFO",
"message": "permit renewal accepted",
"permit_id": "PP-40118",
"duration_ms": 163,
"service": "permit-renewal",
"namespace": "parking",
"pod": "permit-renewal-7c4f9b6d84-x2ktn",
"node": "node-14",
"original_line": "2026-03-11T09:14:27.481Z INFO permit renewal accepted id=PP-40118 ms=163"
}go deeper
Be ready to name the stages in order and say what each does in one sentence: get the bytes, turn them into fields, add context, choose where they go, keep them searchable.
An interviewer expects the mechanics: who tails the file, what re-joins a stack trace, why the read offset matters, and what a half-matching parse pattern does to a subset of lines.
Show you have operated one. Say which stage loses data permanently, what counters you put at each boundary, and how you would replay a day of logs after fixing a bad pattern.
Own the irreversibility argument: which stages the platform makes replayable, what the raw text costs to keep, and how you stop every team inventing its own collection path.
A log aggregation pipeline is everything between a process writing a line and an engineer finding it. It is worth walking it as five stages, because each stage makes a decision and only some of those decisions can be revisited later. Take a municipal parking-permit service: every renewal writes one line. Six months later a regulator asks for evidence that permit `PP-40118` was renewed on a particular morning. Whether that line is findable was settled by five stages, most of them long before anyone knew the question would be asked. ## The five stages | Stage | What it does | The decision it makes | Where it usually runs | | --- | --- | --- | --- | | Collect | reads bytes out of the producing process | which lines enter the pipeline at all | node-level agent, sidecar, in-process | | Parse | turns text into named fields | what becomes a field and what stays opaque text | agent, central tier, or the store at read time | | Enrich | attaches context the emitter did not have | what the record says about where it ran | node-level agent or central tier | | Route | chooses destinations, copies and drops | who can ever see this record | agent or central tier | | Store | indexes, retains, answers queries | what a query can filter on | the backend | **Collect.** On a container platform the process writes to standard output, the runtime writes those bytes to a file on the node, and a node-level log agent — Fluent Bit, Filebeat and Vector are common choices — tails that file and follows rotation. Two things happen here that no later stage repairs cleanly: a stack trace arrives as dozens of physical lines and has to be re-joined into one event, and a very long line may have been written out as fragments that the runtime marks as partial and that must be reassembled before anything parses them. Collect also owns the **read offset** — the pipeline's only record of how far it has got. **Parse.** If the application already emits structured records, parsing is a decode and almost nothing can go wrong. If it emits prose, parsing means matching a pattern — typically a regex, often with named captures. A pattern that fails leaves the line as an opaque blob; a pattern that half-matches is worse, because it silently produces wrong fields for a subset of lines and nobody notices until a query returns too few rows. **Enrich.** The process knows its own request context: the permit id, the trace id, the outcome. It does not know the node it landed on, the namespace, the workload that owns it, the image tag, the cluster or the region. Enrichment is a **join** between the record and something else — the orchestrator's view of the workload, a static configuration, a lookup table. **Route.** Routing is fan-out (the same record to a cheap archive and to a searchable store), splitting by severity or tenant, and dropping classes nobody reads. It is also an access-control decision: a record delivered to a shared destination is visible to everyone who can query that destination, and no later filter can retract it. **Store.** The store accepts records, indexes whatever it indexes, and answers queries. What a query can filter on is bounded entirely by what parse and enrich produced. ## What no later stage can undo - **A line the collect stage never read.** If the agent was stopped, or fell behind while the runtime rotated and deleted the file, that line exists in no other copy. This is the irreversible one. - **Events joined wrongly.** A stack trace split into forty records can be stitched back at query time only clumsily, and a partial fragment dropped on its own leaves a truncated message that reads as complete. - **A field parsed wrongly** is re-derivable *only* if the original text survived alongside it. - **Enrichment inputs are point-in-time.** The facts a record needs — which workload owned this container — may not be resolvable an hour later, so enrichment cannot honestly be deferred to read time. - **A routing drop** is final for the destination that never received it, even if another destination did. - **Store-stage choices are the most recoverable**, because if the raw records still exist upstream you can re-ship them into a different store. ## Designing so that mistakes are survivable 1. **Keep the raw line** next to the parsed fields. It converts a whole class of permanent parse defects into a reprocessing job. 2. **Enrich as early as the facts exist**, and parse as late as you can afford — enrichment inputs expire, patterns do not. 3. **Instrument every boundary**: lines read, records emitted, records acknowledged by the destination, records rejected. A pipeline with no counters cannot tell a quiet service from a broken collector. 4. **Make replay possible.** If the pipeline can be pointed at yesterday's raw data and re-run, every stage after collect becomes a bug you can fix rather than an incident you apologise for. The engineer who can name the five stages is halfway there. The one who can say which stage's mistakes are permanent is the one who designs ingestion that survives its own defects.
- If the store keeps the raw text beside the parsed fields, which stages' mistakes become recoverable and which do not?Parse becomes recoverable: the pattern can be fixed and the stored raw text reprocessed. Enrich becomes only partly recoverable, because the facts it joined against — which workload owned that container — may no longer be resolvable. Collect stays unrecoverable: raw text you never read cannot be re-read. Routing is recoverable only for destinations that still hold a copy.
- Which counters would you put on a log pipeline before you trust it?One per boundary: lines read by the collector, records emitted after parse, records the destination acknowledged, and records rejected or dropped with a reason. The useful signal is the difference between adjacent counters, plus the collector's read lag behind the end of the file. Without those, a pipeline that has silently stopped looks exactly like a quiet service.
- Where does multi-line handling belong, and why not later?As early as possible — in the collect stage, where the physical lines are still adjacent and in order. Once records are batched, routed and possibly reordered by retries, the forty lines of a stack trace are no longer reliably contiguous, so joining them downstream depends on guesswork about ordering and timestamps that the collector never had to make.
It is a postal chain: once a letter is never picked up from the box it is gone, but a wrongly sorted or badly addressed letter can still be re-sorted as long as the envelope survives.
saying these in an interview costs you the question
- Thinks the store can recover a line the collector never read
- Treats parsing and enrichment as one stage
- Assumes every stage runs on the same host
- Believes routing is a setting on the store rather than a pipeline decision
- Cannot name a single place where records are counted or dropped
- Says stack traces are handled by the search backend