What does structured logging give a log consumer that a formatted message string cannot?
answer
- Two shapes for the same event
- Who has to guess the boundaries
- Exact match, not accidental substring
- Comparisons need a real number
- Absence leaves no mark in prose
basics
~20 sStructured logging emits each line as named, typed fields instead of one formatted sentence. Ingest no longer has to guess where each value starts and ends, and queries can filter, compare and aggregate on a field rather than searching raw text.
solid answer
~40 sA printf-style line such as `lunch order 88142 in region eu-north-3 took 412ms` carries the information only as text, so every consumer recovers the values from the sentence's shape — and a reworded message silently breaks that. A structured line emits `order_id`, `region` and `duration_ms` as named values with their own types. Three things follow. **Exact matching** replaces accidental substring matching, so a filter on `region` cannot also hit a line that merely mentions the region. **Typed comparison** becomes possible: `duration_ms > 320` is arithmetic, not string matching. And **aggregation** becomes possible at all — count by `region`, percentile over `duration_ms` — because the store knows which value is which. The cost is that the emitting side now owns a field vocabulary and has to keep it consistent.
code
text · 2 lines2026-02-11T09:14:07Z WARN lunch order 88142 in region eu-north-3 took 412ms
{"timestamp":"2026-02-11T09:14:07.318Z","level":"WARN","service":"order-api","order_id":88142,"region":"eu-north-3","duration_ms":412,"message":"order confirmation exceeded budget"}go deeper
Be ready to show the same event written both ways and name what each field is for. Knowing that a filter on a named field is exact where a text search is not is the core of the answer at this level.
Explain the mechanics: which comparisons need a type, why absence is testable in one shape and invisible in the other, and why a reworded sentence breaks a downstream consumer without failing anything at build time.
An interviewer expects you to weigh the costs you take on — a field vocabulary that is now a contract, higher pre-compression volume, and the fact that nothing enforces consistency across services — against what the query side gains.
Own the question of who governs the field vocabulary across teams, how a rename is rolled out without orphaning saved queries, and what you standardise centrally versus leave to each service.
## The same event, two shapes A log statement records that something happened. In the printf tradition it records it as one formatted sentence: `2026-02-11T09:14:07Z WARN lunch order 88142 in region eu-north-3 exceeded budget, took 412ms` Everything a machine needs is in there, but only as text. In structured form the same event is a set of named values: a timestamp, a severity, a `service`, an `order_id` of `88142`, a `region` of `eu-north-3`, a `duration_ms` of `412`, and a short message. Nothing about the event changed. What changed is that each value now has a **name** and a **type**, and both survive the trip to storage. ## What the ingest side stops guessing Whoever consumes the prose line has to recover `order_id`, `region` and `duration_ms` from the sentence's shape — from the fact that the number after the word "order" is an order id and the token after "took" is a duration in milliseconds. That recovery is a maintained artefact owned by whoever runs the log-processing pipeline, and it carries three standing weaknesses: - **It is coupled to wording.** Someone rephrases the message in a code review, the sentence still reads fine to a human, and the extraction quietly stops matching. Nothing fails at build time and nothing fails at emit time. - **It is ambiguous at the edges.** A value that can itself contain a space, a quote or a newline — a user-supplied name, an exception message, a URL with a query string — breaks the positional assumptions the extraction leans on. - **It is paid late and per line.** The work happens downstream, for every line, forever, instead of once in the code that already knew what each value meant. Structured emission removes the guess: the boundaries were decided by the only party with certain knowledge of them. ## What the query side gains This is the half that decides whether the change is worth anything, and it is not "logs look nicer". Capabilities appear that a substring search over text cannot offer at all. | You want to | Over prose lines | Over named fields | |---|---|---| | Isolate one region | substring match, which also hits lines merely mentioning it | exact match on `region` | | Find everything over the budget | not expressible — the number is characters | `duration_ms > 320` | | Rank regions by their slowest requests | needs the field re-derived on every scanned line | group by `region`, percentile over `duration_ms` | | Find lines that never got an order id | not expressible — absence leaves no mark | a field-is-absent predicate | The second and third rows are the ones to internalise. A numeric comparison needs a number, and a percentile needs a numeric column to compute over; a store that only ever saw characters cannot give you either without re-deriving the value at query time, on every line it touches. The fourth row matters just as much and gets noticed later: in prose, "the field was never set" and "the field was set to an empty value" look identical, because neither leaves a mark a query can test. ## A concrete case A school-meal ordering service runs against a 320 ms p99 budget for the order-confirmation call. A new region is brought up and, for two weeks, ships nothing but prose lines. The question "is the new region inside budget?" is answerable in the established regions in one query and, in the new one, not answerable at all — not slowly, not approximately, but not at all, because no query engine can compute a percentile over a token it has no reason to believe is a number. The gap is not that the new region logs *less*; it logs the same events. It logs them in a shape that cannot be asked questions. ## The costs, stated honestly - **A field vocabulary becomes a contract.** Once `duration_ms` is queried, renaming it to `latency_ms` in one service breaks every saved query, alert and chart built on the old name. - **Volume rises.** Field names repeat on every line. Compression absorbs much of it, but ingest billing is usually metered before compression. - **Raw lines get harder to read.** An object tailed in a terminal is worse than a sentence. The fix is rendering — emit fields in production, pretty-print them locally — not a return to prose. - **Nothing enforces consistency.** The emitter is now free to name and type values however it likes, which is a real and common failure mode rather than a theoretical one. ## The rule of thumb Anything you would ever filter, sort, group or compare on is a field. Anything that exists purely so a human reading one line understands what happened stays in the message. Interpolating a value into the message *and* emitting it as a field is normal and is not duplication worth removing: the field serves the query engine, the message serves the person, and they have different requirements.
- If the values are already fields, why keep a human-readable message at all?Because a person still reads one line at a time during an incident and needs the point of it without decoding a dozen fields. A constant message text is also the cheapest grouping key for "how often does this event class happen". Keep the sentence constant and put the variable parts in fields; interpolating an id into the sentence produces a million distinct texts and destroys that grouping.
- Does structured logging have to mean JSON?No. JSON is the common encoding, but a key/value line, a tab-separated record with a declared header, or any binary encoding the pipeline understands all qualify. What defines structured logging is that each value arrives with a name and a type decided by the emitter, not that any particular syntax was used to carry it.
- What is the practical cost of renaming a field once it is in use?It breaks every saved query, alert condition and chart built on the old name, and it splits historical data: old lines carry the old name, new lines the new one, so any query spanning the change has to match both. Treat a queried field name as a published contract and dual-emit through a transition rather than swapping it in one deploy.
A prose log is a paragraph about a receipt; a structured log is the receipt's line items. You can total the second one.
saying these in an interview costs you the question
- Says structured logging just means writing logs in JSON
- Treats a substring search as equivalent to an exact field match
- Thinks numeric comparison works without a real numeric type
- Claims the human-readable message should be deleted once fields exist
- Assumes renaming a queried field is a free internal change
- Interpolates ids into the message text, destroying grouping