skip to content

In structured logs, one service emits `attempt` as a number and another as a string. What breaks?

level: seniorimportance: nice to knowfreq 26%

answer

  1. Same field, two services, two types
  2. The first record decides
  3. Refused, coerced, or dropped
  4. The emitter never hears about it
  5. A hole looks like good news

basics

~20 s

A store that fixes each field's type when it first sees it refuses the conflicting records or drops that field. The emitting service sees no error, so the drifting lines simply go missing from queries.

solid answer

~50 s

Drift comes in three shapes and they fail differently. **Type conflict** — `attempt` numeric here, `"3"` there — is the sharp one: a store that fixed the type at first sight has already committed earlier data, so the later record is refused outright or has that field dropped. **Depth drift**, where one service nests a value three objects deep, produces a different queryable path for the same meaning and multiplies distinct field names. **Presence drift**, a field emitted on most lines and omitted on some, means those lines quietly fail to match any filter on it. All three are silent because the loss happens past a one-way boundary: the emitter got no error, so no exception, no retry and no health signal moves. What you see is a hole in a result set — and a hole looks like good news.

code

json · 2 lines
json
{"service":"order-api","region":"eu-west-1","attempt":2,"message":"supplier lookup retried"}
{"service":"order-api","region":"eu-north-3","attempt":"2","message":"supplier lookup retried"}

go deeper

for a junior

Be ready to say that a structured log field has a type as well as a name, and that two services sending the same field with different types is a problem rather than something the store quietly reconciles.

for a middle

Explain what a store that fixes types at write time does when a conflicting value arrives, and why the emitting service never learns about it. Know that an absent field and an empty field are different query outcomes.

for a senior

An interviewer expects you to reason about the shape of the loss — that refused records are a correlated subset, not a random one — and to name the two detection signals: the store's reject count and per-service accepted-line rates.

for a principal

Own the prevention: a shared field vocabulary enforced in code rather than in a document, an explicit nesting budget, and a policy for optional fields so queries are written knowing which fields may be absent.

## Three shapes drift takes "Schema drift" in structured logs is not one problem. It is three, and they fail differently. 1. **Type drift.** `attempt` is a number in one service and the string `"3"` in another. Same name, same meaning, incompatible representation. 2. **Depth drift.** One service emits `user_id` at the top level; another nests it three deep under an object, so the queryable path is different even though nothing about the value changed. 3. **Presence drift.** A field is emitted on most lines and quietly omitted on some — an error path that returns before the value is known, a code path that never set it. All three come from the same root cause: a structured log line's shape is decided independently by every service that writes one, and nothing in the emitting process validates it against what anyone else is emitting. ## What a schema-on-write store does with each A store that fixes a field's type when it first sees it has no good options once a conflicting value arrives, because the earlier data is already committed under the earlier type. | Drift | Typical outcome | Who notices | |---|---|---| | Type conflict | the whole record is refused, or that field is dropped and the rest kept | nobody, unless someone watches the reject count | | Deep nesting | a distinct queryable path per depth; the number of distinct field names grows | whoever pays for the store, eventually | | Sometimes absent | the line simply does not match any filter on that field | whoever trusts a filtered count | Two of those three are silent by construction, and the third is silent for weeks. ## Why it is silent The failure happens on the far side of a one-way boundary. The emitting service handed a line to a local writer and got no error, because from its point of view the write succeeded — the loss occurs later, in a component that has no channel back to the code that produced the data. So there is no exception, no failed request, no retry, and nothing in the service's own health signals moves. What you see instead is a *hole*: a query returns fewer results than it should, a dashboard panel that used to have a line on it goes flat, a count that everyone trusts is quietly wrong. A hole looks exactly like good news. "That service stopped erroring" and "that service's error lines stopped being accepted" render identically. ## A worked case A school-meal ordering service runs a 320 ms p99 budget on order confirmation. A new region is brought up in a hurry with no instrumentation of its own, and its ordering service — built from a slightly older template — emits `attempt` as a string where the established regions emit it as a number. The store fixed `attempt` as numeric months ago. Roughly 12.4% of the new region's lines carry `attempt`, and every one of those records is refused. The dashboards do not break. They show the new region running comfortably inside budget, because the lines that carried retries — the slow ones, the ones with `attempt` above 1 — are precisely the ones being thrown away. The reported p99 is a percentile over a biased sample, and it is biased in the flattering direction. This is the reason type drift is worth an interview question at all: it does not degrade the data uniformly, it removes a correlated subset. ## Catching it before it costs you - **Own the field vocabulary in one place.** A shared logging helper or a small library that both names and types the common fields is the only intervention that scales past a handful of services; a written convention with no code behind it drifts within two quarters. - **Alert on rejects, not on volume alone.** The store knows how many records it refused and why. If nobody has wired that count to anything, drift is undetectable by design. - **Watch per-service line rates.** A service whose accepted-line count falls sharply while its request rate is flat is the signature of drift, and it is visible without touching the log content. - **Budget nesting explicitly.** Flatten to one or two levels and use a compound field name rather than depth; deep, free-form objects turn every distinct sub-key into another queryable field name. - **Decide the absent case per field.** Either the field is always present with an explicit "unknown" value, or it is genuinely optional and every query on it is written knowing that. What hurts is an optional field that everyone treats as mandatory. ## The mental model Structured logs are a schema you did not write down. Type drift is what happens when a schema exists implicitly, is enforced somewhere far from the code, and reports its violations to nobody. The fix is never "be more careful in code review" — it is to make the shape a shared artefact and to put a monitored signal on the enforcement point.

  • Why is losing the drifting lines worse than losing a random sample of the same size?
    Because the loss is correlated with the thing you are measuring. If `attempt` is only present on retried requests, refusing those records removes the slow tail specifically, and the surviving percentile is biased in the flattering direction. A uniformly random loss of the same volume would leave the distribution's shape intact; a correlated loss silently rewrites it.
  • How would you detect this without inspecting individual log lines?
    Watch two signals. First, the store's own rejected-record count with its reason — it knows exactly what it refused, and if nothing is wired to that number the failure is undetectable by design. Second, accepted lines per service against that service's request rate: a sharp fall in accepted lines while traffic is flat is the signature, and it needs no access to log content.
  • What actually prevents drift across dozens of services?
    A shared logging helper that both names and types the common fields, so the vocabulary lives in code rather than in a convention document. Written conventions with nothing enforcing them drift within a couple of quarters as services are forked from stale templates. Pair it with a monitored reject count so the remaining escapes surface quickly.

saying these in an interview costs you the question

  • Assumes any store accepts both types under one field name
  • Expects the emitting service to receive an error on conflict
  • Believes a missing field reads as zero or empty in queries
  • Nests log objects arbitrarily deep with no field budget
  • Reads a drop in accepted line volume as a traffic change
  • Thinks code review alone is enough to keep field types aligned