skip to content

Which fields should every structured log line carry, whatever the logging stack?

level: middleimportance: must knowfreq 68%

answer

  1. What must a single line answer
  2. When, how bad, who, which work
  3. Ordering across hosts needs an offset
  4. One name is not one process
  5. Humans still read lines one at a time

basics

~20 s

Every line needs a timestamp with an explicit offset and stated precision, a severity level, an identity for the emitting service and instance, the identifiers that make the line joinable to a unit of work, and a short human-readable message.

solid answer

~50 s

The set is short because it is driven by three uses: filtering a line out of millions, sorting it beside lines from other machines, and reading it alone. So: a **timestamp** with an explicit offset — normally UTC in the ISO 8601 / RFC 3339 shape — at millisecond or finer precision, because events inside one request share a second. A **severity level** as its own field, not a word inside the sentence. A **service identity plus an instance identity**, because a service name cannot isolate one sick replica out of forty. The **joinable identifiers** — request id, trace id, job run id, tenant or order id — spelled identically in every service, or the join gets written twice and wrong once. And a **constant message text**, so a person can read one line and so lines group by event class.

code

json · 12 lines
json
{
  "timestamp": "2026-02-11T09:14:07.318Z",
  "level": "WARN",
  "service": "order-api",
  "instance": "order-api-7c4f9d-x2ktb",
  "env": "prod",
  "region": "eu-north-3",
  "request_id": "6f1c9a4e-2b70-4d31-9a18-c0e5f7412d63",
  "order_id": 88142,
  "duration_ms": 412,
  "message": "order confirmation exceeded budget"
}

go deeper

for a junior

Be ready to list what a single line must answer: when, how bad, who emitted it, which request it belongs to, and what happened. Knowing that each of those is a named field rather than part of a sentence is enough here.

for a middle

An interviewer expects the mechanics: why an explicit offset is required for ordering across hosts, what precision buys you inside one request, and why a service name without an instance hides a single failing replica.

for a senior

Show the operational judgement — keeping emitter time separate from arrival time, keeping message text constant so event classes group, and treating a field name as a contract that several teams must spell the same way.

for a principal

Own the standard itself: which fields are mandatory fleet-wide, how they are enforced in shared code rather than in a document, and how you migrate a field name once dashboards and alerts across many teams depend on it.

## What "every line" is supposed to mean A structured log line is only useful if it can be found and understood without its neighbours. In practice a line is read in one of three ways: pulled out of a million others by a filter, sorted next to lines from other machines, or stared at alone after a filter narrowed things down. Those three uses are what determine the mandatory field set, and they are the reason the set is short. ## The five things a line has to answer 1. **When did it happen?** A timestamp, with an explicit timezone offset and a stated precision. 2. **How bad is it?** A severity level, as its own field. 3. **Who emitted it?** A service identity, plus an identity for the specific running instance. 4. **What unit of work does it belong to?** The identifiers that make the line joinable — a request id, a trace id, a job run id, a tenant or order id. 5. **What happened?** A short human-readable message. Everything else — the values specific to this one event — is added on top, and the useful discipline is that those extras are also fields rather than fragments of the sentence. ## Timestamps: the field most often got wrong Three separate mistakes live here. - **No offset.** A wall-clock time with no zone is unorderable the moment two hosts sit in different zones, and unrecoverable afterwards, because nothing in the line says what it was relative to. Emit a full date-time with an explicit offset — the ISO 8601 / RFC 3339 shape, normally in UTC. - **Unstated or insufficient precision.** Events inside one request routinely land in the same millisecond. Whole-second precision makes the ordering of a request's own lines a coin flip. - **Confusing event time with ingest time.** The store also stamps a line when it arrives. Those two diverge under buffering, retries and backfill, sometimes by hours. Keep the emitter's own timestamp as a field of its own; the arrival stamp is the pipeline's business, not the event's. ## Identity: a service name is not enough `service` answers "which program" and is what you filter on nine times out of ten. It cannot answer "which of the forty replicas", so a single sick host — one bad disk, one stale configuration, one process that missed a restart — hides inside the aggregate. Carry both: the logical service and the concrete instance, whatever your platform calls the latter. The same argument runs one level out: the deployment environment and the region belong on the line too, because "works in one region, fails in another" is a question you will ask, and a filter you cannot build after the fact. | Field group | Answers | Failure if omitted | |---|---|---| | Timestamp with offset | when | lines from different zones cannot be ordered | | Severity | how bad | urgency has to be inferred from wording | | Service + instance | who | one sick replica is invisible inside the fleet | | Joinable identifiers | which unit of work | one request's story cannot be reassembled | | Message | what happened | a line is unreadable without a query engine | ## The identifiers that make a line joinable A single line is nearly always a fragment. The value of the line is that it can be put back with the other fragments of the same request, job or order, and that only works if every one of them carries the same identifier under the same name. Two disciplines follow. First, the identifier is generated once, at the edge of the system, not per-component. Second, it is spelled identically everywhere: `trace_id` in one service and `traceId` in another means the join has to be written twice and will be written wrong once. ## Why the human-readable message survives It is tempting to conclude that a fully structured line no longer needs prose. It does, for reasons that are practical rather than sentimental. - **Someone reads one line at a time.** During an incident, a person scanning results needs the point of the line without decoding a dozen fields. - **A stable message is a grouping key.** Lines that share a message text are the same event class, which is how you count "how often does this happen" cheaply. - **It is where the un-fielded context goes.** Not every nuance justifies a field; a sentence absorbs the rest without inventing vocabulary. The rule that keeps it honest: the message is a **constant string**, and the variable parts are fields. `"order confirmation exceeded budget"` with `order_id` and `duration_ms` alongside it groups cleanly; interpolating the id into the sentence produces a million distinct message texts and destroys grouping. Interpolating a value that is *also* a field is fine for readability, as long as the grouping-critical part of the sentence stays constant.

  • The store already stamps every line on arrival. Why keep the emitter's own timestamp?
    Because the two diverge, sometimes by hours. Buffering, retries and backfill all delay arrival, so ordering by the arrival stamp reorders events that happened in a fixed sequence. The emitter's timestamp is the only record of when the thing actually happened; the arrival stamp answers a different question — when the pipeline saw it — and both are worth keeping for exactly that reason.
  • Why does severity need to be a field when the message often makes the seriousness obvious?
    Because severity is filtered and ordered on more than any other field, and a query engine cannot rank English adjectives. As a field it supports "everything at or above warning in the last hour" as one cheap predicate; as prose it requires matching an open-ended set of phrasings that changes every time someone rewords a line.
  • How much precision does the timestamp actually need?
    Enough to order the lines of a single request, which in practice means milliseconds at minimum and microseconds where a request emits many lines in a tight loop. Whole-second precision makes the ordering within a request arbitrary, and since reconstructing a sequence is the main reason to look at one request's lines, that is a real loss rather than a cosmetic one.

saying these in an interview costs you the question

  • Logs local wall-clock time with no timezone offset
  • Emits a service name but no per-instance identity
  • Puts the severity level only inside the message text
  • Assumes the store's arrival stamp can replace event time
  • Interpolates ids into the message so nothing groups
  • Spells the same identifier differently in each service