skip to content

How would you structure Elasticsearch component and index templates for many teams and index families?

level: principalimportance: should knowfreq 26%

answer

  1. few semantic layers, one owner each
  2. published number ranges prevent collisions
  3. JSON in git, applied by a pipeline
  4. assert on the composed result, not the fragments
  5. a change reaches data only at the next index

basics

~20 s

Layer a small number of semantic component templates - platform defaults, shared schema, per-team fields, an override slot - and give each family one narrow index template with a priority from a documented band. Keep all of it in version control and gate changes on simulated output.

solid answer

~50 s

Three decisions carry the design. **Layering**: a handful of component templates with clear ownership — platform-wide settings, a shared field schema, per-domain fields, and a reserved high-precedence override component — assembled by one index template per family. Deep chains make "where did this setting come from?" unanswerable. **Priority bands**: publish reserved ranges (platform baseline, team families, incident overrides) and keep every band above the priorities of Elastic's shipped templates for patterns you claim, so teams can add templates without colliding and without silently losing to a vendor template. **Change control**: templates live in git, are applied by CI, carry `version` and `_meta` for provenance, and every merge diffs `_simulate_index` output for the real index names against golden files. Then plan for latency: a template change only reaches data at the next index creation, so the rollout plan for each family is a rollover or a reindex, and that is a capacity and scheduling decision, not an afterthought.

code

bash · 15 lines
bash
# layers, each with one owner, listed in precedence order
PUT _index_template/orders-events
{
  "index_patterns": ["orders-events-*"],
  "priority": 320,
  "composed_of": [
    "platform-defaults",
    "schema-common",
    "orders-fields",
    "orders-events-overrides"
  ],
  "ignore_missing_component_templates": ["orders-events-overrides"],
  "version": 7,
  "_meta": { "owner": "orders-team", "source": "infra/es-templates/orders-events.json" }
}

go deeper

for a junior

Not a junior question. Focus on being able to read an existing template and say which component contributes what before worrying about how a fleet of them is organised.

for a middle

You can contribute by keeping your team's component narrow, listing composed_of in true precedence order and simulating before you open the pull request.

for a senior

Be ready to design one family's templates end to end and to run the cutover: pick the priority, verify with simulate, and schedule the rollover or reindex that makes the change real.

for a principal

Own the conventions themselves - layer ownership, published priority bands, the CI gate on composed output, and the rollout calendar - and be able to justify each against the failure it prevents at fleet scale.

## What actually goes wrong at scale A single team with three index families needs no governance. The failure modes appear at a dozen teams and a hundred families, and they are consistent: - Two teams pick the same priority for overlapping patterns and the second `PUT` fails, or worse, they pick patterns that overlap a vendor template and their configuration silently never applies. - A shared component template gains a field for one team and it lands in everyone's indices. - Nobody can answer "why does this index have five shards?" because it takes reading six component templates and knowing merge order. - A mapping fix is deployed and takes effect three weeks later when a family finally rolls over, in an order nobody tracked. Governance is the set of conventions that make each of these either impossible or loud. ## Layer by ownership, not by convenience Give component templates semantic roles with a single owner each. A workable default: 1. **platform-defaults** — cluster-wide index settings the platform team owns: replica count, refresh interval, the ILM policy reference. Nobody else edits it. 2. **schema-common** — the field schema every document shares (`@timestamp`, service and environment identifiers, correlation ids). Changes here are cross-cutting and reviewed as such. 3. **\<domain\>-fields** — one per team or bounded context, owned by that team, holding only its own fields. 4. **\<family\>-overrides** — a reserved, highest-precedence slot per family, usually empty, that exists so an incident fix has somewhere to go without editing shared layers. Each family then gets exactly one index template with narrow `index_patterns` and `composed_of: [platform-defaults, schema-common, domain-fields, family-overrides]`. Precedence is positional, so this list reads exactly as the precedence order, which is the property that makes it reviewable. From 8.7 the override slot can be referenced with `ignore_missing_component_templates` so the index template stays valid while the component does not exist — the same hook Elastic's own integrations leave for user customisation. ## Publish the priority bands `priority` is a bare integer with no built-in meaning, so it needs an external convention. Reserve ranges and write them down: - 0–99: nothing (leave room below the vendor templates rather than fighting them) - 100–199: vendor and integration-managed templates — do not occupy - 200–299: platform baseline templates covering broad families - 300–499: team family templates - 500+: deliberate overrides, including incident response The absolute numbers matter less than that they are documented, that per-family templates sit above per-family-group ones, and that you have checked the priorities of any shipped templates covering your patterns with `GET _index_template`. Ban catch-all `*` patterns outright: a high-priority `*` template claims every index in the cluster including ones you have never heard of. ## Treat templates as code with a real gate Every template is a JSON file in a repository, applied by a pipeline, never by hand in a console. Populate `_meta` with owner, ticket and source path, and bump `version` on each change, so an operator staring at `GET _index_template` can attribute anything they find. The gate that earns its keep is the composed-output diff. In CI, apply the templates to a disposable cluster, run `POST _index_template/_simulate_index/<name>` for every index name pattern the system genuinely creates, and diff the resulting settings and mappings against checked-in golden files. This catches the whole silent class: a dropped field, a `keyword` flipped to `text`, a component reorder, a new integration template that now outranks yours. It also makes review honest, because the artefact reviewers read is the effective mapping rather than a fragment whose merged meaning nobody can compute by eye. ## Plan the latency, not just the change The hardest part to explain to stakeholders is that a merged template change reaches no data at all until an index is created. Every rollout therefore has a second half: which families roll over, when, and which need a reindex because they are not rolling. Reindexing a large family is real I/O and real cluster capacity, so the rollout is scheduled work. For additive, non-breaking field changes you can simply wait for natural rollover; for a type correction, the family must be reindexed behind an alias or accept a mixed mapping across generations — and mixed mappings across an aliased set are their own operational hazard when queries span both. ## Guard the blast radius Two limits belong in the design. Keep an eye on total field count per index — a shared schema component that grows unchecked pushes families toward the field limit and, more importantly, toward slow, memory-hungry mappings. And keep the number of component layers small: every layer added is a permanent tax on every future debugging session. When someone proposes a fifth layer, the right question is whether the concern is genuinely orthogonal to the existing four or is simply a team-shaped fragment that belongs in that team's own component.

  • How do you stop a shared schema component from becoming a dumping ground?
    Give it an owner and a rule: only fields genuinely common to every document belong there, and anything team-specific goes in that team's own component. Enforce it in review and watch total field count per index, since an unbounded shared schema pushes every family toward slow, memory-hungry mappings.
  • How do you coexist with vendor or integration-managed templates on the same index patterns?
    Read their priorities with GET _index_template and reserve a band strictly above them for anything you must override, leaving their range unoccupied. Prefer the documented customisation hook — an override component the integration already references — over outranking their template wholesale, which forfeits future updates.
  • What is the rollout plan once a template change merges?
    Additive field changes can wait for natural rollover. Type corrections cannot: the family needs a reindex behind an alias, or you accept different mappings across generations, which breaks queries that span them. Schedule the reindexes as capacity work and track which families have actually picked the change up.

saying these in an interview costs you the question

  • Assigns template priorities ad hoc with no published bands
  • Uses a catch-all * pattern at high priority
  • Builds deep component chains nobody can trace
  • Reviews template fragments instead of the composed output
  • Treats a merged template change as if it were already live

context