What does each conventional log level from TRACE to FATAL mean, and what test separates WARN from ERROR?
answer
- A contract with the on-call reader
- Ask who has to act
- Recovered, or not recovered
- WARN handled it; ERROR did not
- TRACE and DEBUG are developer detail
basics
~20 sLog levels are a contract about who must act. ERROR means something failed and a human should look; WARN means something recovered but degraded; INFO records state changes; DEBUG and TRACE are developer detail, off by default in production.
solid answer
~40 sTreat the ladder as a contract with whoever reads the line during an incident. **ERROR** means a unit of work failed and nothing recovered it — someone has to look. **WARN** means the system detected something wrong and handled it: a retry succeeded, a fallback fired, a default was applied. The separating test is *actionability*: if no human action would change the outcome, it is not an ERROR. **INFO** records the state changes a later reader needs to reconstruct what the process did — startup, configuration resolved, a batch finishing. **DEBUG** is the detail needed to follow one code path; **TRACE** is every branch and payload. **FATAL**, where an ecosystem has it, means the process is terminating. The corollary catches most teams out: a rejected malformed request is the service working, not an ERROR.
code
text · 3 lines14:02:11.418 WARN inventory.reconcile vault-7 humidity read timed out, retry 1/3 succeeded after 812ms
14:02:14.907 ERROR inventory.reconcile vault-7 humidity read failed after 3 retries; batch 4471 abandoned
14:02:14.908 INFO inventory.reconcile batch 4471 finished: 1183 lots written, 1 vault skippedgo deeper
Be ready to name the ladder in order and say in one sentence what each level is for. The one an interviewer will push on is WARN versus ERROR, so have a concrete example of each from code you have actually written.
An interviewer at this level expects the actionability test rather than a list. Explain why a handled-and-recovered condition is not an ERROR, and how mislabelled severity makes an alerting stream worthless once readers learn to ignore it.
Show that you set severity for the person on call. Talk about auditing what a service emits at ERROR, pushing a noisy dependency's lines down with a per-logger threshold, and alerting on rates rather than on individual occurrences.
Own severity as a fleet-wide convention: without an agreed test for each level, cross-service dashboards and routing built on severity measure nothing. Be ready to say how you would get dozens of teams to agree and how you would detect drift afterwards.
## What a log level actually communicates A log level is not a measure of how interesting the author found a line. It is a claim about **who has to do something about it, and how soon**. Everything downstream is built on that claim: which lines reach an alerting stream, which are routed somewhere low-priority, which are shed at the edge before they cost anything, and above all what an on-call engineer's eye skips over at three in the morning. A ladder set by mood instead of by contract quietly disables all of it. The conventional ladder, most to least severe, is `FATAL` (absent from some ecosystems), `ERROR`, `WARN`, `INFO`, `DEBUG`, `TRACE`. A threshold of INFO means INFO and everything above it is emitted, while DEBUG and TRACE are stopped by the level check and cost almost nothing. Other ladders exist — syslog's severities include `Alert` and `Notice`, which most application frameworks do not have — so "the six levels" is a convention, not a universal. | Level | The claim being made | Who acts | |---|---|---| | FATAL | The process cannot continue and is terminating | Whoever owns the service, immediately | | ERROR | A unit of work failed and nothing recovered it | A human, during this shift | | WARN | Something was wrong and the system handled it | Eventually, or whoever watches the trend | | INFO | The system changed state in a way worth recording | Nobody; it is context for the next reader | | DEBUG | The detail needed to follow one code path | The engineer who turned it on | | TRACE | Essentially everything: branches, payloads, loops | The engineer who turned it on, briefly | ## The test that separates WARN from ERROR Ask one question of the line: **did anything recover the work?** - A timeout that a retry then satisfied — **WARN**. The caller got its answer. - A primary dependency that failed while a cached fallback served the request — **WARN**, and a metric too, because the fallback rate is what actually matters. - A configuration key absent and a documented default applied — **WARN**, once, at startup. - A request that failed after every retry and returned a server error — **ERROR**. Nothing saved it. - A background job that abandoned a batch — **ERROR**, because the work simply did not happen. The corollary catches most teams out: **a rejected bad request is not an ERROR.** A client sending a malformed payload and getting a rejection is the service working exactly as designed. Logging it at ERROR is the most common way a service ends up with an ERROR stream everyone has learned to ignore — at which point the level carries no information and the first real failure goes unnoticed. Log it low, keep enough identity to find it later, and let a rate metric carry the trend. The discipline runs in the other direction too. If a library you depend on logs at ERROR for a condition your code handles cleanly, lower that specific logger's threshold rather than letting its judgement pollute yours. Severity is a claim about *your* system, and the library author cannot see your recovery path. ## What DEBUG and TRACE are for Both are developer instruments rather than reader instruments, and what separates them is granularity, not audience: 1. **DEBUG** answers "which path did this take?" — the branch chosen, the parameters resolved, the decision made. It is what you turn on for one package when chasing a specific behaviour. 2. **TRACE** answers "what happened at every step?" — loop iterations, intermediate values, payload contents. It exists precisely because it is far too much to have on by default. Neither is written for whoever is paged. That is why production normally runs at INFO: not because DEBUG is forbidden, but because a line nobody will read still costs argument formatting, allocation and throughput on every request that produces it. ## Where the contract breaks in practice Across a 41-service estate the failure is never that the level names are unknown; it is that the test behind them is applied differently in every service. One team's ERROR is another's WARN, and any dashboard, routing rule or on-call filter built on severity across those services is then measuring nothing coherent. During one incident on a cheese-ageing inventory platform, the humidity-reconciliation service had been logging the failing condition roughly 4,000 times an hour for six weeks — at WARN, because a retry usually rescued it — while the alerting stream watched only ERROR. Nothing was hidden. The severity had simply made a claim ("handled, nobody need act") that stopped being true when the retries began failing too, and no one revisited it. Two habits prevent that. First, choose the level for the reader you expect rather than for the code you happen to be in: ask who would act on the line. Second, treat severity drift as a reviewable defect — audit periodically what a service emits at ERROR, be willing to move lines down, and be equally willing to promote a WARN whose recovery has stopped being reliable. A level ladder is worth exactly as much as the consistency with which a team applies it.
- Where does a validation failure caused by a bad client request belong — ERROR or something lower?Below ERROR. A rejected request is the system working: the caller sent something invalid and was told so. Logging it at ERROR trains everyone that ERROR is noise and buries the failures that mean the service is broken. Log it at INFO or DEBUG with enough identity to find it later, and track the rate as a metric so you alert on the trend rather than on individual lines.
- A library you depend on logs at ERROR for conditions your service handles fine. What do you do?Fix it at the boundary rather than absorbing the noise. Lower that specific logger's threshold so the library's ERROR lines stop reaching the alerting stream, and log your own line at the level that reflects your handling. Per-logger thresholds exist for exactly this: severity is a claim about your system, and the library author cannot see your recovery path.
- What is the difference between FATAL and ERROR, and why do some ecosystems not have FATAL?FATAL means the process cannot continue and is terminating; ERROR means one unit of work failed while the process keeps serving. Some ecosystems omit FATAL because the distinction is rarely used well — a dying process usually leaves an ERROR plus a crash signal anyway — and a level nobody sets consistently is worse than no level at all.
A level is a triage tag, not a volume knob: it says who has to act on the line, not how interesting the author found it.
saying these in an interview costs you the question
- Says ERROR just means something bad happened
- Logs every caught exception at ERROR regardless of recovery
- Treats levels as verbosity for the author, not urgency for the reader
- Claims WARN and ERROR are interchangeable in practice
- Thinks INFO is the right level for per-request debug detail
- Assumes every ecosystem ships the same six levels