skip to content

Several teams ship logs into one Elastic Stack cluster. How do you make a common field schema stick, and what breaks when two services send one field name with different types?

level: principalimportance: should knowfreq 41%

answer

  1. Field names become an interface
  2. Enforce on the path, not in a wiki
  3. One index rejects; many indices go quiet
  4. Keep free-form values in one string namespace
  5. A new meaning gets a new name

basics

~20 s

Agree a common vocabulary such as the Elastic Common Schema and enforce it in the shared shipping tier, not in documentation. When two services disagree on a field's type, one index rejects documents and searches across several go quietly wrong.

solid answer

~40 s

Field names are an interface once several teams share a store, so adopt a published vocabulary — the Elastic Common Schema, which Beats and Elastic Agent already emit — and mandate a small core: `@timestamp`, `message`, `log.level`, `service.name`, `event.dataset`, `trace.id`. Enforce it where producers cannot route around it, by renaming into the common names in Logstash filters or an ingest pipeline, and validate the rest in each service's build. A type disagreement fails two ways: inside one index the type is fixed by the first document and later documents of the other type are rejected; across indices searched together nothing fails at all, but ranges, sorts and numeric aggregations over that field stop meaning anything. The second is worse because it is silent, and the only real repair for months of history is re-indexing.

code

text · 12 lines
text
filter {
  mutate {
    rename => {
      "svc"        => "[service][name]"
      "lvl"        => "[log][level]"
      "elapsed_ns" => "[event][duration]"
    }
  }
  mutate {
    convert => { "[event][duration]" => "integer" }
  }
}

go deeper

for a junior

Recall that logs from different services land in one place and that everyone using the same field names is what makes searching across them possible. Know that a field also has a type, not just a name.

for a middle

Explain what a shared vocabulary such as the Elastic Common Schema fixes, name a few of its core fields, and describe how a producer's own field names are renamed into it on the way in.

for a senior

Show the operational consequence: which failure is loud, which is silent, how you detect a type conflict early through rejections and dead-letter growth, and why re-indexing is the only real repair for history.

for a principal

Own the tradeoff. Decide how much schema is mandatory versus optional, whether the cost lands on the platform tier or on producing teams, and hold the line that a shared field never changes meaning — a new meaning gets a new name.

## What a shared field vocabulary actually buys Once several teams' logs land in one store, field names stop being a local styling choice and become an interface. Elastic publishes one such vocabulary — the Elastic Common Schema — and Beats and Elastic Agent already emit it, which is why adopting it is usually cheaper than inventing your own. It fixes both the *name* and the *type* of common concepts: `@timestamp`, `message`, `log.level`, `service.name`, `service.version`, `host.name`, `event.dataset`, `error.message`, `url.path`, `http.response.status_code`, `source.ip`, `user.name`, `trace.id`, `span.id`. Four things become possible only when everyone agrees: - **One query works everywhere.** "Errors for this service in the last hour" is a single expression instead of one per team's spelling of `level`, `lvl`, `severity` and `loglevel`. - **Dashboards and alerts become reusable assets.** A platform team can ship one error-rate panel that works for a service it has never heard of, on the day that service goes live. - **Correlation with traces is free.** If every service writes `trace.id` and `span.id` with those names, moving from a log line to the request it belongs to is a link, not a project. - **Onboarding stops being bespoke.** The cost of adding the fortieth service becomes the cost of adding the second. The corresponding discipline is that ECS is large, and enforcing all of it is a good way to get the whole idea rejected. Pick a core subset that every producer must emit and leave the rest optional. ## What happens when two services disagree about a type Say the billing service writes `duration` as a JSON number of milliseconds and the meter-reading service writes it as the string `"310ms"`. Two distinct failures follow, and confusing them is the usual weak answer. | Where the two land | What you observe | |---|---| | The same index | The type is fixed by the first document seen. Documents of the other type are rejected outright, which surfaces as indexing failures — and, if the shipping tier has a dead-letter queue, as a growing pile there. | | Different indices searched together | Both index fine. The field simply has two types across the set, so range filters, sorts and numeric aggregations over it stop being meaningful. | The second case is the dangerous one, because nothing fails. Data keeps arriving, dashboards keep rendering, and the defect is only discovered when someone asks a question that spans both. A field that has been ambiguous for four months cannot be fixed by changing the producer — the historical documents still carry the wrong type, and the only real repair is re-indexing the affected data, which on an 18 GB-a-day stream is a multi-day operation with a real cluster-load cost. That is the shape of the argument to bring to an interview. On a district-heating billing platform, a regulator asking for six months of evidence over a field that was a number for two of those months and a string for four is not a query problem; it is a data problem that was created by an absent contract and can now only be paid for. ## Making the contract stick Documentation does not make a schema stick — the shipping path does. The levers, roughly in order of how reliably they hold: 1. **Normalise in the shared tier.** Rename each producer's fields into the common names in Logstash filters or an ingest pipeline attached to the index. This is the only mechanism that holds when a team ignores the standard, because the team cannot route around it. 2. **Emit the schema from a shared logging library.** Cheaper to run and gives correct data at source, but any service that opts out of the library opts out of the contract. 3. **Validate before production.** A schema check in each service's build that fails when a log field is not in the agreed vocabulary catches drift while it is still free to fix. 4. **Constrain the free-form space.** Put anything genuinely per-service into a single namespace whose values are all strings — ECS's `labels` object is designed for exactly this — so a new key can never introduce a new type into the shared vocabulary. 5. **Alert on the symptom.** Watch indexing rejections and dead-letter-queue growth; a type conflict announces itself there long before anyone notices a broken dashboard. The genuinely principal judgement is *how much* to enforce and *where the cost lands*. Central normalisation is a permanent tax on the platform team and grows one rule per producer; producer-side emission is cheaper to run but needs organisational authority the platform team may not have. The usual settlement is a small mandatory core enforced centrally at the shipping tier, a much larger optional vocabulary that teams adopt because the dashboards are already written for it, and a hard rule that a shared field never changes type — a new meaning gets a new name. Renaming a field is cheap. Redefining one is a re-index.

  • A field has been ambiguous for four months. What are your actual options?
    Fixing the producer stops the bleeding but repairs nothing already stored. The options are re-indexing the affected history into a corrected shape, which costs cluster time proportional to the volume; querying each side separately and reconciling outside the store; or accepting the gap and documenting it. Choose deliberately and say so out loud — quietly leaving a broken field in place is how the same question comes back a year later.
  • How do you let teams add their own fields without reopening the type problem?
    Give them one namespace whose values are all strings — ECS's `labels` object is built for it — so an arbitrary new key can never introduce a new type into the shared vocabulary. Anything that needs to be a number, a date or a boolean has to be proposed as a named field and reviewed. That keeps the free-form space genuinely free while keeping the shared space stable.
  • Why enforce normalisation in the shipping tier rather than in a shared logging library?
    Because the library is opt-in and the tier is not. A library gives correct data at source and costs the platform team nothing to run, which makes it the better default; but any service that does not use it, or pins an old version, silently leaves the contract. The tier is the backstop that holds regardless, at the price of one rename rule per producer.

saying these in an interview costs you the question

  • Treating a field-naming standard as documentation rather than enforcement
  • Assuming a type conflict always shows up as a visible error
  • Believing fixing the producer repairs already-stored history
  • Mandating an entire large schema instead of a core subset
  • Letting arbitrary keys land as arbitrary types in shared fields
  • Reusing a field name for a new meaning instead of adding one