skip to content

How do you govern a structured log-record schema across services, and what does it cost?

level: principalimportance: should knowfreq 30%

answer

  1. It is an interface, not a preference
  2. Small fixed core, namespaced extras
  3. Future interpreter names can collide
  4. Machine time, not host-local time
  5. Somebody has to own and version it

basics

~20 s

Treat the record shape as a data contract: a small stable core of fields, application fields namespaced under one key, machine-friendly UTC timestamps, an explicit rule on secrets, and additive-only changes. The costs are bytes, serialization CPU and human readability.

solid answer

~50 s

The decision is not "JSON or text" but "who consumes this, and what are they promised". Fix a small core every service emits — timestamp, level, logger name, message, correlation id, service and version — and put everything application-specific under a namespaced key so a future CPython `LogRecord` attribute cannot collide with it and raise `KeyError` at every call site, as `taskName` did in 3.12. Make timestamps deterministic: emit `record.created` as an epoch float, and if you also render text, set `Formatter.converter = time.gmtime` with a pinned format rather than a locale-sensitive code, or the same code emits different strings on different hosts. Ship one internal formatter so the contract has an owner and a test. Price it honestly — serialization runs per handler per record and bytes are billed downstream — and keep a human-readable handler for developers. Changes are additive first, with a schema version field.

code

python · 15 lines
python
import logging, sys, time


class UtcFormatter(logging.Formatter):
    converter = time.gmtime
    default_time_format = "%Y-%m-%dT%H:%M:%S"
    default_msec_format = "%s.%03dZ"


handler = logging.StreamHandler(sys.stdout)
handler.setFormatter(UtcFormatter("%(asctime)s %(levelname)s %(message)s"))
log = logging.getLogger("thumbnails")
log.addHandler(handler)
log.setLevel(logging.INFO)
log.info("nightly batch started")

go deeper

for a junior

Know that logs are read by tools as well as people, and that a field such as a job id is far easier to search for than the same value glued into a sentence.

for a middle

Be ready to say which fields belong in every record and why timestamps should be UTC and machine-parseable rather than formatted for the host you happen to be on.

for a senior

Show that you can implement the contract: one formatter, a namespaced field area, an allow-list for what gets emitted, and a rule that keeps secrets and unbounded-cardinality values out of the stream.

for a principal

Own it as an interface. Justify the core field set, the cost per record in CPU and bytes, the migration path for a rename, and the boundary between what logs answer and what metrics and traces answer better.

## What you are actually deciding Once logs are parsed by machines, the record shape stops being a formatting preference and becomes a published interface. A dashboard, an alert rule, a retention policy and an on-call runbook all bind to field names. Renaming `job_id` to `jobId` is then an API break with no compiler to catch it. So the governance question — what is in the record, who may add to it, and how a change ships — matters more than the choice of serializer. ## A small core plus a namespaced extension area The core is the set of fields every consumer may assume: a timestamp, a level, the logger name, the message text, a correlation id, the service name and its deployed version. Keep it short; a field in the core is a promise every service must keep forever. Everything else goes under one namespaced key — an `app` or `ctx` object — rather than at the top level. This buys two things. It stops teams colliding with each other, and it protects you from CPython itself: keys passed via `extra` are written straight onto the `LogRecord`, and `Logger.makeRecord` raises `KeyError` when the name is already taken. The set of taken names grows between releases — `taskName` was added in 3.12 and immediately reserved that word — so a top-level field name is a bet on future interpreter versions that a namespaced one does not make. Interpreter upgrades should not be able to turn a log call into an exception. ## Determinism beats prettiness A log record read by a machine wants values it cannot misread. Emit the timestamp as `record.created`, an epoch float: no timezone, no locale, sortable, exact. If you also want human-readable text, pin it — `Formatter.converter = time.gmtime` with an explicit ISO-like `default_time_format` — because `Formatter.formatTime` otherwise runs `time.localtime` through `time.strftime`, and codes such as `%c` or `%x` render according to the host's locale settings. That is not theoretical. A nightly six-hour thumbnail batch is the kind of job that gets moved onto a spare host, and if the format string was left locale-dependent the same code emits a differently-shaped timestamp there — so the pipeline's parser rejects the batch's records, and you discover it when you go looking for the run that failed and find no logs at all. Any field a human might localize — timestamps, decimal separators, month names, sorted enum labels — belongs in an invariant machine form, with localization left to whatever displays it. ## Price it Structured logging is not free and a lead should be able to say what it costs. - **CPU**: serialization runs once per handler per record, on the calling thread. Records that never pass a level check cost nothing, which makes level discipline the biggest lever you have; where building the *arguments* is expensive, `logger.isEnabledFor` is the guard. - **Bytes**: JSON is several times the size of the equivalent line, and downstream storage and ingest are usually billed by volume. Compact separators, short key names and dropping empty fields are real savings at scale. - **Cardinality**: a field with unbounded distinct values is cheap to emit and expensive to index. Ids belong in fields; free-form user text does not. - **Humans**: JSON is miserable to read in a terminal. Keep a plain-text handler for local development and interactive debugging — formatters are per handler precisely so both can coexist. ## Safety rules that belong in the contract Structured fields are indexed, retained and readable by more people than the service's own team, so a secret in a field is worse than the same secret in prose. State the rule where the schema is defined: no credentials, tokens or personal data as fields; log identifiers rather than values; cap field sizes; prefer an explicit allow-list of emitted fields over dumping the record's whole attribute dictionary, so a hasty `extra={"user": user}` cannot export an object's entire repr into a searchable store. Where user-controlled text does get logged, it is data, not markup — serializing through `json.dumps` escapes it, which is one more reason to build the payload as a dict rather than concatenating strings. ## Ownership and change Give the contract a home: one internal formatter and filter, wired the same way in every service, with tests that assert the core fields exist and parse. Then changes have a place to happen. Additions are safe and go first; a removal or rename runs on a clock — emit both shapes, migrate the consumers, then drop the old one — and a `schema` version field makes it possible for a consumer to know which shape it has. This is the same discipline as any other published interface, and the fact that a log line looks informal is exactly why teams forget to apply it. ## When not to do it A command-line tool, a one-off script and a local development run are all better served by readable text. Structured records earn their cost when something other than a human reads them; if nothing does, JSON is ceremony that makes debugging harder.

  • What is your policy for fields whose values come from user-controlled data?
    An allow-list, a size cap and no secrets or personal data — log an identifier instead of the value. Build the payload as a dict and let `json.dumps` escape it, so the text stays data rather than becoming structure. And watch cardinality: a field with unbounded distinct values is cheap to write and expensive to index, which is a cost the logging team pays and the emitting team never sees.
  • How do you roll out a change to the record schema without breaking consumers?
    Additively. New fields first, consumers updated, then the old field removed on an announced date; a `schema` version field lets a consumer tell the shapes apart. Because a new top-level field name can also collide with a future `LogRecord` attribute, additions land inside the namespaced object rather than at the root — which makes the change safe against both the consumers and the interpreter.
  • When would you argue against structured JSON logging?
    When nothing but a human reads the output: command-line tools, local development, short scripts. JSON costs bytes, CPU and readability, and it buys correlation and queryability — if no collector is consuming it, you have paid the price for none of the benefit. The usual answer in a service is both: a JSON handler for the pipeline and a plain-text handler for a terminal.

saying these in an interview costs you the question

  • Says JSON logging is free at any volume
  • Dumps the whole record dictionary and lets the pipeline sort it
  • Leaves timestamps in host local time or a locale format
  • Treats the schema as whatever each service happens to emit
  • Claims structured logs replace metrics and tracing
  • Adds top-level field names without reserving against future attributes

context