Where in a log pipeline should parsing into fields happen — the emitting process, a node-level agent, a central processing tier, or at query time?
answer
- Four places, four different bills
- Who pays the CPU, who owns the change
- Structured at source has no pattern
- A wrong rule at the edge is widest
- Query time defers the cost to readers
basics
~20 sParsing at the source is cheapest and has no pattern to break, but needs every application changed. A node agent spreads the CPU and rolls out slowly. A central tier fixes rules in one place. Query time bills every reader.
solid answer
~50 sThere is no single right answer, only four different bills. **In the emitting process**: the application writes structured records, so no pattern can break and no CPU re-derives what the code already knew — but every service must change, and much of the volume is code you do not own. **In a node-level agent**: CPU spreads across the fleet, but a rule change is a fleet rollout and a wrong rule damages records where the original text is least likely to survive. **In a central tier**: one place to fix a rule and replay it, one place to overload, and a blast radius covering every team. **At query time**: nothing is decided at ingest, and every reader pays again. Large estates end up hybrid: structured at source where they own the code, agent-side parsing for stable high-volume formats, and a central tier for the awkward ones.
code
text · 1 line^(?<ts>\S+) (?<level>[A-Z]+) (?<msg>.+?) id=(?<permit_id>\S+) ms=(?<duration_ms>\d+)$go deeper
Know that a text log has to be turned into fields before anyone can filter on them, and that an application which logs structured records directly saves the pipeline from guessing.
Be able to compare parsing at the source, in an agent and in a central service in terms of CPU and of how quickly a rule can be changed when a format shifts.
Show you have debugged a bad pattern in production: silent partial matches, greedy captures, backtracking that stalls a node, and the parse-failure metric that would have caught it.
Own it as policy for a fleet: a default placement, an exception path for formats you do not control, an explicit stance on retaining raw text, and who is accountable when a rule is wrong.
Parsing means turning a line of text into named fields — timestamp, severity, message, and whatever else the line carries. The question is not *how* to parse but *where*, and it is a genuine architecture decision because the four options differ in who pays the CPU, who can change the rules, and what happens when the rules are wrong. Run the example: a municipal parking-permit service emits `2026-03-11T09:14:27.481Z INFO permit renewal accepted id=PP-40118 ms=163`. Somewhere between that process and a search box, `permit_id` and `duration_ms` have to become fields. ## The four placements | Placement | CPU cost | Flexibility | Blast radius of a wrong rule | | --- | --- | --- | --- | | Emitting process | lowest — the code already has the values | lowest — a change is a code change and a deploy | one service, caught in its own tests | | Node-level agent | spread thinly over every node | medium — a config rollout across the fleet | every workload on every node using that rule | | Central processing tier | concentrated, must be capacity-planned | highest — one change, effective immediately | every team at once, but replayable | | Query time, in the store | paid again on every query, by every reader | highest — change your mind per query | none at ingest; every reader repeats the mistake | **In the emitting process.** The application logs structured records directly. Nothing downstream has to guess what a field is, because nothing was ever flattened into prose. This is the cheapest option in total CPU and the only one with no pattern to break. Its limits are ownership: you cannot change a database engine's log format, a reverse proxy's access log, or a vendor image, and in a large estate those produce a large share of the volume. **In a node-level agent.** The agent tailing the container log files parses as it reads. The CPU cost is genuinely distributed — a hundred nodes each doing a hundredth of the work — and the parsed record is smaller on the wire than the text it replaced. The cost is change latency and blast radius: updating a pattern means rolling a configuration change across every node, and until it lands, records from the affected services are wrong. Worse, the mistake happens at the point where the original text is most likely to be discarded. **In a central processing tier.** All records arrive raw at a service that parses and forwards. One pattern change fixes the whole estate immediately, and if the tier keeps or receives the raw text, records can be replayed after a fix. The costs are real: the tier must be sized for peak ingest and becomes a shared failure domain — a pattern that makes one team's regex catastrophically slow degrades everyone's ingest, not just theirs. **At query time.** Nothing is extracted at ingest; readers apply an extraction expression when they search. This is maximally flexible — a question you did not anticipate is still answerable a month later — and it removes ingest-time data loss entirely. The bill arrives on every query, repeatedly, and knowledge of the format moves into the heads of readers and into saved queries instead of living in one place. ## What "wrong pattern" actually costs A regex that fails outright is the friendly case: the line stays opaque and someone notices. The dangerous cases are quieter: - **A partial match** extracts fields for the 94% of lines that fit and silently drops the rest, so a dashboard undercounts and looks healthy. - **A greedy pattern** absorbs a field into the message, so filtering by that field returns nothing while the data appears to be present. - **A catastrophically backtracking pattern** turns one unusual line into seconds of CPU. On a node agent this stalls collection for every workload on that node; in a central tier it stalls the estate. - **A parse that runs before multi-line assembly** treats each line of a stack trace as its own event. ## How to decide 1. **Split by ownership, not by taste.** Services you own emit structured records at source; everything else is parsed by the pipeline. 2. **Push parsing towards the source for high-volume, stable formats** — an access log with a fixed shape is a good node-agent job. 3. **Keep the awkward, changing formats central**, where a fix is one change and the raw text is still around to reprocess. 4. **Leave the long tail to query time.** Fields that one team wants once a quarter do not deserve ingest-time cost or an ingest-time risk. 5. **Preserve the raw text past the first parse wherever the cost is bearable.** It is what converts a wrong-pattern incident into a reprocessing job. 6. **Instrument parse failure as a first-class rate**, per rule, and alert on a change in it. The absence of that signal is why bad patterns survive for months. The judgement an interviewer is listening for is that this is not one choice for the estate but a policy: a default placement, an explicit exception path, and a way to notice when a rule is quietly wrong.
- You cannot change a vendor image that emits prose logs. Where do you parse it and why?In the pipeline, not the process, because the code is not yours to change. Prefer a central tier over node agents for a format you do not control: vendor formats change on upgrade, and a central rule can be corrected in one place and replayed against retained raw text, whereas a node-level rule needs a fleet rollout while records are being written wrong.
- What signal tells you a parse rule broke, without anyone opening a dashboard?A per-rule parse-failure rate, exported as a metric from wherever parsing runs, alerted on relative change rather than an absolute threshold. Pair it with a records-in versus fields-extracted ratio per source. A rule that starts failing after a deploy shows up as a step change within a scrape or two, long before a human notices that a query returns fewer rows than it should.
- Does emitting structured records at source remove the need for a pipeline parse stage?It removes the pattern, not the stage. Records still need decoding and validating, and the pipeline still has to cope with a service that emits malformed output, a startup banner written as prose before the logger initialises, and crash output written by the runtime rather than the application. Those never arrive structured, whatever the application's own logging does.
saying these in an interview costs you the question
- Says parsing always belongs in the agent, with no tradeoff named
- Ignores that most volume comes from code you do not own
- Forgets a wrong rule can silently drop a subset of lines
- Treats query-time extraction as free because ingest is cheaper
- Never mentions who can change the rule or how fast
- Assumes one placement must serve the whole estate