skip to content

How would you design mappings for a multi-tenant Elasticsearch log platform where tenants send arbitrary JSON keys?

level: principalimportance: should knowfreq 30%

answer

  1. Two zones, not one schema
  2. Someone else writes the fields you must live with
  3. Settings bound growth; something else bounds blast radius
  4. The cheapest field never reaches the cluster
  5. Rolled-over indices let mistakes expire

basics

~20 s

Split the document into a declared core that is strictly mapped and a free-form subtree that is not allowed to mint unbounded fields. Isolate tenants across indices so one tenant's key space cannot poison a shared mapping, and enforce field-count budgets with alerting.

solid answer

~50 s

I would treat the document as two zones. The **core** — timestamp, tenant, service, level, message, trace ids — is explicitly mapped in an index template with `dynamic: strict` at the root, so a typo or a rogue field fails loudly at the edge rather than quietly enlarging the schema. The **tail** lives in one designated subtree whose `dynamic` is `false` or `runtime`, or which is typed for arbitrary key spaces, so tenant keys stay retrievable and queryable without becoming permanent mapping entries. Dynamic templates enforce house rules on anything that is allowed through — detected strings as `keyword`, no full-text analysis on machine fields. Then **isolation**: tenants get their own data streams so a mapping problem has a blast radius of one, with `index.mapping.total_fields.limit` as a per-index guardrail plus alerting well below it. Finally, time-based indices mean today's mistake ages out instead of living forever, and an ingest pipeline normalises names before Elasticsearch ever sees them.

code

json · 22 lines
json
{
  "mappings": {
    "dynamic": "strict",
    "dynamic_templates": [
      { "labels_as_keyword": {
          "path_match": "labels.*",
          "match_mapping_type": "string",
          "mapping": { "type": "keyword", "ignore_above": 256 }
      }}
    ],
    "properties": {
      "@timestamp": { "type": "date" },
      "tenant":     { "type": "keyword" },
      "service":    { "type": "keyword" },
      "level":      { "type": "keyword" },
      "message":    { "type": "text" },
      "labels":     { "type": "object", "dynamic": "runtime" },
      "raw":        { "type": "object", "dynamic": "false" }
    }
  },
  "settings": { "index.mapping.total_fields.limit": 500 }
}

go deeper

for a junior

Focus on the mechanism rather than the strategy: unknown keys can be rejected, ignored, or made query-time only, and unbounded keys are what break an index.

for a middle

Explain how a strict core plus a constrained free-form subtree is assembled — templates, dynamic settings on an object, dynamic templates for house rules — and what each choice costs.

for a senior

Argue the operational side: per-tenant isolation for blast radius, field-count alerting below the limit, an ingest pipeline that normalises before Elasticsearch sees the data, and a runbook for when the alert fires.

for a principal

Own the cost transfers — strictness pushes work to producers, runtime fields push it to query time, isolation pushes it to shard count — and state which you are buying, what measurement would change your mind, and how the contract is published and governed.

## Framing the decision The question is not "which setting" but "where does the schema contract live". Elasticsearch will happily accept whatever arrives and turn it into permanent, cluster-wide metadata. In a multi-tenant platform the people producing the fields are not the people operating the cluster, so the design's job is to make the cost of a new field visible and bounded before it lands. ## Two zones in every document **A declared core.** Every log line has a small set of fields the platform itself depends on: `@timestamp`, tenant identifier, service, log level, message, correlation ids. These are queried by every dashboard, they must be typed consistently across tenants, and they are worth full index-time cost. Map them explicitly in an index template and set `dynamic: strict` at the mapping root. Strictness here is a feature: a producer that renames `service` to `svc` gets an immediate rejection instead of silently disappearing from every dashboard. **A bounded tail.** Everything tenant-specific lands under one subtree — `labels`, `attributes`, whatever you name it — and that subtree is where you spend the design effort. The options, from cheapest to most capable: - `dynamic: false` — values stay in `_source` and are visible in the document view, but nothing is searchable until someone defines a field for it. Zero mapping growth. - `dynamic: runtime` — new keys become runtime fields: discoverable and queryable without index-time cost, at the price of per-document evaluation on queries that touch them. - A field type designed for arbitrary keys, which indexes a whole object under a single mapping entry, trading some query expressiveness for a fixed schema footprint. - Restructuring into key/value pairs at ingest: `[{"k": ..., "v": ...}]` gives two fields no matter how many keys exist, at the cost of pair-matching semantics unless you make it nested. There is no single winner. The right choice depends on whether tenants need to *search* their custom fields (runtime or key/value) or merely *see* them (`false`), and how often. ## Isolation is the real control Settings bound growth; isolation bounds blast radius. A shared index means every tenant's key space contributes to one mapping, so the worst-behaved tenant sets the failure date for everyone, and you cannot fix one tenant without reindexing all of them. Per-tenant data streams cost more indices and more shards — which is its own scaling limit — so the realistic design is tiered: large or high-variance tenants get dedicated streams, the long tail of small tenants shares a pooled stream with a strict tail policy, and there is a documented trigger for promoting a tenant out of the pool. ## Guardrails and observability `index.mapping.total_fields.limit` should be treated as a **circuit breaker, not a budget**. The budget is an alert at a fraction of it, so a tenant's new field pattern becomes a ticket days before it becomes a write outage. In time-based data the strongest signal is the delta: compare the field count of today's backing index with yesterday's, and investigate any jump. Companion caps — depth and nested-field limits — deserve the same treatment. When the alert fires, the runbook should already exist: identify the offending key pattern, decide whether it is legitimate, and either tighten the subtree, add a dynamic template, or move the tenant. Raising the limit is an explicit, approved exception with an owner, not a reflex. ## Push the contract upstream The cheapest field is the one that never arrives. An ingest pipeline can rename, drop or fold unrecognised keys before mapping happens, which turns "an incident later" into "a rule now". Where the platform has an SDK, the same normalisation belongs there, with the server-side pipeline as the enforcement layer for producers who bypass it. Publishing the contract — these fields are indexed and queryable, these are stored only, this is how you request a new indexed field — makes the tenant a participant rather than a hazard. ## Time as an eraser Time-based indices give you something a single long-lived index never has: a schedule on which mistakes expire. Fix the template today, and every index created afterwards is correct; retention removes the damaged ones without any reindex at all. That is why arbitrary-key data belongs in rolled-over indices even when volume alone would not require it. ## The trade you are actually making Every axis here is a cost transfer. Strict mapping moves cost from the cluster to the producers, who now have to declare fields. Runtime fields move cost from ingest and storage to query time, which is fine until many users query the same tail field. Per-tenant isolation moves cost from mapping size to shard count, and shard count has its own ceiling. A good answer names which of these you are choosing to pay, why it fits this platform's traffic shape, and what measurement would make you revisit the choice.

  • Why give large tenants their own data streams instead of tuning one shared index harder?
    Because a shared mapping has a shared failure. One tenant's key explosion consumes everyone's field budget, and remediation means reindexing data belonging to tenants who did nothing wrong. Separate streams give each tenant its own mapping and its own limit, so the blast radius is one customer — paid for in extra indices and shards, which is why only the large or high-variance tenants get them.
  • How do you decide between dynamic: runtime and simply storing the tail with dynamic: false?
    By whether tenants must search those fields. `runtime` keeps them queryable and discoverable with no index-time cost, but every query touching one runs a script per examined document, so it degrades as usage grows. `false` costs nothing at all and keeps values visible in the document, but they are unsearchable until someone defines a field. Start restrictive and promote fields that earn it.
  • What would make you revisit this design a year in?
    Query telemetry showing a handful of tail fields dominating traffic — that is the signal to promote them into the declared core. Rising per-index field counts despite the guardrails means the contract is being bypassed upstream. Growing shard counts from tenant isolation mean the pooling threshold needs raising. Each has a measurement, and each points at a different lever.

saying these in an interview costs you the question

  • Proposes one shared index for all tenants with a big field limit
  • Treats the field limit as a budget to fill rather than a breaker
  • Ignores who produces the fields and where the contract is enforced
  • Makes everything a runtime field and calls it free
  • Assumes a bad mapping can be cleaned up without reindexing

context