skip to content

Logging Concepts

How production logging works as a system — formats, levels, pipelines, storage models, and cost. Interviewers probe this because logging is every team's biggest observability bill and noisiest signal.

on this pageshow

questions

19

What does each conventional log level from TRACE to FATAL mean, and what test separates WARN from ERROR?

level: juniorimportance: must knowfreq 66%

answer

  1. A contract with the on-call reader
  2. Ask who has to act
  3. Recovered, or not recovered
  4. WARN handled it; ERROR did not
  5. TRACE and DEBUG are developer detail

basics

~20 s

Log levels are a contract about who must act. ERROR means something failed and a human should look; WARN means something recovered but degraded; INFO records state changes; DEBUG and TRACE are developer detail, off by default in production.

solid answer

~40 s

Treat the ladder as a contract with whoever reads the line during an incident. **ERROR** means a unit of work failed and nothing recovered it — someone has to look. **WARN** means the system detected something wrong and handled it: a retry succeeded, a fallback fired, a default was applied. The separating test is *actionability*: if no human action would change the outcome, it is not an ERROR. **INFO** records the state changes a later reader needs to reconstruct what the process did — startup, configuration resolved, a batch finishing. **DEBUG** is the detail needed to follow one code path; **TRACE** is every branch and payload. **FATAL**, where an ecosystem has it, means the process is terminating. The corollary catches most teams out: a rejected malformed request is the service working, not an ERROR.

code

text · 3 lines
text
14:02:11.418 WARN  inventory.reconcile  vault-7 humidity read timed out, retry 1/3 succeeded after 812ms
14:02:14.907 ERROR inventory.reconcile  vault-7 humidity read failed after 3 retries; batch 4471 abandoned
14:02:14.908 INFO  inventory.reconcile  batch 4471 finished: 1183 lots written, 1 vault skipped

go deeper

for a junior

Be ready to name the ladder in order and say in one sentence what each level is for. The one an interviewer will push on is WARN versus ERROR, so have a concrete example of each from code you have actually written.

for a middle

An interviewer at this level expects the actionability test rather than a list. Explain why a handled-and-recovered condition is not an ERROR, and how mislabelled severity makes an alerting stream worthless once readers learn to ignore it.

for a senior

Show that you set severity for the person on call. Talk about auditing what a service emits at ERROR, pushing a noisy dependency's lines down with a per-logger threshold, and alerting on rates rather than on individual occurrences.

for a principal

Own severity as a fleet-wide convention: without an agreed test for each level, cross-service dashboards and routing built on severity measure nothing. Be ready to say how you would get dozens of teams to agree and how you would detect drift afterwards.

## What a log level actually communicates A log level is not a measure of how interesting the author found a line. It is a claim about **who has to do something about it, and how soon**. Everything downstream is built on that claim: which lines reach an alerting stream, which are routed somewhere low-priority, which are shed at the edge before they cost anything, and above all what an on-call engineer's eye skips over at three in the morning. A ladder set by mood instead of by contract quietly disables all of it. The conventional ladder, most to least severe, is `FATAL` (absent from some ecosystems), `ERROR`, `WARN`, `INFO`, `DEBUG`, `TRACE`. A threshold of INFO means INFO and everything above it is emitted, while DEBUG and TRACE are stopped by the level check and cost almost nothing. Other ladders exist — syslog's severities include `Alert` and `Notice`, which most application frameworks do not have — so "the six levels" is a convention, not a universal. | Level | The claim being made | Who acts | |---|---|---| | FATAL | The process cannot continue and is terminating | Whoever owns the service, immediately | | ERROR | A unit of work failed and nothing recovered it | A human, during this shift | | WARN | Something was wrong and the system handled it | Eventually, or whoever watches the trend | | INFO | The system changed state in a way worth recording | Nobody; it is context for the next reader | | DEBUG | The detail needed to follow one code path | The engineer who turned it on | | TRACE | Essentially everything: branches, payloads, loops | The engineer who turned it on, briefly | ## The test that separates WARN from ERROR Ask one question of the line: **did anything recover the work?** - A timeout that a retry then satisfied — **WARN**. The caller got its answer. - A primary dependency that failed while a cached fallback served the request — **WARN**, and a metric too, because the fallback rate is what actually matters. - A configuration key absent and a documented default applied — **WARN**, once, at startup. - A request that failed after every retry and returned a server error — **ERROR**. Nothing saved it. - A background job that abandoned a batch — **ERROR**, because the work simply did not happen. The corollary catches most teams out: **a rejected bad request is not an ERROR.** A client sending a malformed payload and getting a rejection is the service working exactly as designed. Logging it at ERROR is the most common way a service ends up with an ERROR stream everyone has learned to ignore — at which point the level carries no information and the first real failure goes unnoticed. Log it low, keep enough identity to find it later, and let a rate metric carry the trend. The discipline runs in the other direction too. If a library you depend on logs at ERROR for a condition your code handles cleanly, lower that specific logger's threshold rather than letting its judgement pollute yours. Severity is a claim about *your* system, and the library author cannot see your recovery path. ## What DEBUG and TRACE are for Both are developer instruments rather than reader instruments, and what separates them is granularity, not audience: 1. **DEBUG** answers "which path did this take?" — the branch chosen, the parameters resolved, the decision made. It is what you turn on for one package when chasing a specific behaviour. 2. **TRACE** answers "what happened at every step?" — loop iterations, intermediate values, payload contents. It exists precisely because it is far too much to have on by default. Neither is written for whoever is paged. That is why production normally runs at INFO: not because DEBUG is forbidden, but because a line nobody will read still costs argument formatting, allocation and throughput on every request that produces it. ## Where the contract breaks in practice Across a 41-service estate the failure is never that the level names are unknown; it is that the test behind them is applied differently in every service. One team's ERROR is another's WARN, and any dashboard, routing rule or on-call filter built on severity across those services is then measuring nothing coherent. During one incident on a cheese-ageing inventory platform, the humidity-reconciliation service had been logging the failing condition roughly 4,000 times an hour for six weeks — at WARN, because a retry usually rescued it — while the alerting stream watched only ERROR. Nothing was hidden. The severity had simply made a claim ("handled, nobody need act") that stopped being true when the retries began failing too, and no one revisited it. Two habits prevent that. First, choose the level for the reader you expect rather than for the code you happen to be in: ask who would act on the line. Second, treat severity drift as a reviewable defect — audit periodically what a service emits at ERROR, be willing to move lines down, and be equally willing to promote a WARN whose recovery has stopped being reliable. A level ladder is worth exactly as much as the consistency with which a team applies it.

  • Where does a validation failure caused by a bad client request belong — ERROR or something lower?
    Below ERROR. A rejected request is the system working: the caller sent something invalid and was told so. Logging it at ERROR trains everyone that ERROR is noise and buries the failures that mean the service is broken. Log it at INFO or DEBUG with enough identity to find it later, and track the rate as a metric so you alert on the trend rather than on individual lines.
  • A library you depend on logs at ERROR for conditions your service handles fine. What do you do?
    Fix it at the boundary rather than absorbing the noise. Lower that specific logger's threshold so the library's ERROR lines stop reaching the alerting stream, and log your own line at the level that reflects your handling. Per-logger thresholds exist for exactly this: severity is a claim about your system, and the library author cannot see your recovery path.
  • What is the difference between FATAL and ERROR, and why do some ecosystems not have FATAL?
    FATAL means the process cannot continue and is terminating; ERROR means one unit of work failed while the process keeps serving. Some ecosystems omit FATAL because the distinction is rarely used well — a dying process usually leaves an ERROR plus a crash signal anyway — and a level nobody sets consistently is worse than no level at all.

A level is a triage tag, not a volume knob: it says who has to act on the line, not how interesting the author found it.

saying these in an interview costs you the question

  • Says ERROR just means something bad happened
  • Logs every caught exception at ERROR regardless of recovery
  • Treats levels as verbosity for the author, not urgency for the reader
  • Claims WARN and ERROR are interchangeable in practice
  • Thinks INFO is the right level for per-request debug detail
  • Assumes every ecosystem ships the same six levels
open as a page

What does a log backend that full-text indexes every line at ingest buy you, and what does it cost?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A full-text index built at ingest makes any word in any line searchable without declaring it first, and makes aggregations fast. You pay twice: CPU on the write path for every record, and index bytes stored alongside the logs.

open as a page

What does structured logging give a log consumer that a formatted message string cannot?

level: juniorimportance: must knowfreq 76%

basics

~20 s

Structured logging emits each line as named, typed fields instead of one formatted sentence. Ingest no longer has to guess where each value starts and ends, and queries can filter, compare and aggregate on a field rather than searching raw text.

open as a page

In a log aggregation pipeline, what does each stage — collect, parse, enrich, route, store — do, and which decisions can no later stage undo?

level: middleimportance: must knowfreq 70%

basics

~20 s

Collect reads bytes out of the producing process, parse turns text into named fields, enrich attaches context the emitter lacked, route chooses destinations, store indexes and serves. Collect losses are unrecoverable: an unread line exists in no other copy.

open as a page

What drives the cost of a centralized log platform, and which of those does shortening retention actually reduce?

level: middleimportance: must knowfreq 64%

basics

~20 s

Four things drive a log bill: bytes ingested, the parsing and index-building done on them at write time, stored bytes multiplied by replicas and retained days, and query load. Shortening retention shrinks only the third.

open as a page

Which fields should every structured log line carry, whatever the logging stack?

level: middleimportance: must knowfreq 68%

basics

~20 s

Every line needs a timestamp with an explicit offset and stated precision, a severity level, an identity for the emitting service and instance, the identifiers that make the line joinable to a unit of work, and a short human-readable message.

open as a page

How do you turn a service's log level up at runtime without a redeploy, and why per-logger rather than globally?

level: middleimportance: should knowfreq 44%

basics

~20 s

A runtime change has to mutate the live logger objects, not just a configuration file, and invalidate the effective level cached on every descendant logger. Scope it per-logger so one package gets detail while everything else stays at INFO.

open as a page

In a tiered log store, what differs between a hot, a warm and a cold copy of the same data?

level: middleimportance: should knowfreq 47%

basics

~20 s

Tiers differ in the media the bytes sit on, whether the data is still directly searchable and how fast, and how many copies exist. Moving down a tier lowers the price per gigabyte; only deletion removes the bytes.

open as a page

In a log store that indexes only stream labels and scans the lines, which query shapes are fast and which are catastrophic?

level: middleimportance: should knowfreq 52%

basics

~20 s

Only the labels identifying each stream are indexed, so a query selects streams and a time range, then fetches, decompresses and scans those lines. Narrow labels over a short window are fast; a long window with no label matcher is catastrophic.

open as a page

A log shipping agent cannot reach its destination for 40 minutes — what are its options, and what does a disk-backed buffer change over an in-memory one?

level: seniorimportance: should knowfreq 58%

basics

~20 s

It can buffer in memory, buffer to disk, stop reading and let pressure build, or drop. Disk survives an agent restart and holds far more, at the cost of node disk and I/O. Retry after a failed acknowledgement means duplicates.

open as a page

In a log collection pipeline, which records do you discard or sample before ingest, and what makes that irreversible?

level: seniorimportance: should knowfreq 56%

basics

~20 s

Discard or sample the classes with high volume and no demonstrated readers: repetitive health-endpoint access lines, duplicated stack traces, oversized fields. It is irreversible because those records never reach storage, so version the rules and count what you discard.

open as a page

How do you choose between a full-text-indexed log store and a label-indexed scanning one, and what later says you chose wrong?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Decide from the query log, not the product page: if almost every search already knows the service and time window, buy cheap writes and pay per query; if searches are open-ended or aggregate across everything, buy the index.

open as a page

How does a logging context carry per-request fields into a deep log call, and how does it silently break?

level: seniorimportance: should knowfreq 52%

basics

~20 s

A logging context is an ambient, scope-bound key/value map the logging library merges into every record, so no field is passed as an argument. It breaks when work moves to another thread: the map was bound to the thread, not the work.

open as a page

Where in a log pipeline should parsing into fields happen — the emitting process, a node-level agent, a central processing tier, or at query time?

level: principalimportance: should knowfreq 44%

basics

~20 s

Parsing at the source is cheapest and has no pattern to break, but needs every application changed. A node agent spreads the CPU and rolls out slowly. A central tier fixes rules in one place. Query time bills every reader.

open as a page

What does a node-level log agent know about a containerised workload that the process itself does not, and what happens to that metadata when the pod is rescheduled?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

A node-level log agent knows where the process runs — node, namespace, pod, workload — by joining the file it reads against the orchestrator's view of that node. That view is point-in-time: once the pod is deleted, late lines cannot be enriched.

open as a page

What does schema-on-write commit a log store to that schema-on-read does not?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

Schema-on-write freezes the fields, their types and what is searchable before data lands, so mistakes are permanent for records already written. Schema-on-read keeps the raw line and derives structure per query, so the interpretation stays changeable and retroactive.

open as a page

One logger is emitting thousands of lines a second in production. How do you cut that volume without hiding the failure?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

Rate-limit or deduplicate at the call site rather than deleting the line: keep the first occurrence, count the rest, and emit a summary saying how many were suppressed. A silent drop is worse than the noise it removed.

open as a page

In structured logs, one service emits `attempt` as a number and another as a string. What breaks?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

A store that fixes each field's type when it first sees it refuses the conflicting records or drops that field. The emitting service sees no error, so the drifting lines simply go missing from queries.

open as a page