skip to content

When would you use terms_lookup instead of inlining values in a terms query?

level: middleimportance: should knowfreq 40%

answer

  1. The list already lives somewhere as data
  2. Keep the request body small
  3. Server-side fetch of an array from a document
  4. Name an index, an id, and a path
  5. One tiny index copied to every node

basics

~20 s

Use terms_lookup when the list of values already lives in another document — a user's group memberships or blocked ids — so Elasticsearch fetches it server-side instead of your application shipping thousands of values in every request.

solid answer

~50 s

A `terms` query matches documents whose field holds any of the values you list, with no analysis applied. When that list is large and already stored somewhere — a user's followed-author ids, a tenant's permitted categories — inlining it means your app reads it, serializes it, and sends it over the wire on every search. `terms_lookup` instead names an `index`, an `id`, and a `path`, and Elasticsearch fetches the array from that document's `_source` itself. The list stays in one place and requests stay small. The costs are real: it is an extra internal get on every search, the lookup document is read fresh so a stale-cache story does not exist but a hot-shard story does, and the inline form is still capped by `index.max_terms_count`. The usual production shape is a tiny lookup index with one primary shard replicated to every node so the fetch is always local.

code

json · 15 lines
json
{
  "query": {
    "bool": {
      "filter": {
        "terms": {
          "author_id": {
            "index": "user_follows",
            "id": "u-42",
            "path": "following"
          }
        }
      }
    }
  }
}

go deeper

for a junior

Be ready to write a terms query with an inline array and to say that the values are matched exactly, not analyzed.

for a middle

Explain the lookup form field by field — index, id, path — and where the values are read from, and know that the inline list has a configurable ceiling.

for a senior

Show the operational reasoning: a dedicated single-shard lookup index replicated everywhere, awareness of the extra get per search, and when denormalization beats lookup outright.

for a principal

Own the access-control data model: where entitlement lists live, how they propagate, and the cost profile of resolving them per query versus baking them into the documents.

## What the terms query does `{"terms": {"status": ["open", "pending"]}}` matches documents whose `status` field contains at least one of the listed values. It is term-level, so **none of the values are analyzed** — they are looked up byte for byte in the term dictionary. Everything that makes a `term` query fail on a `text` field applies here identically, only multiplied by the length of the list. `terms` is effectively a `bool` of `should` clauses collapsed into one efficient clause, and it does not compute per-term relevance: all matches get the same constant score, which makes it a natural filter-context clause. The inline list is bounded by the index setting `index.max_terms_count`, which defaults to 65,536. Exceeding it fails the request rather than degrading quietly — a useful guardrail, because a request carrying tens of thousands of literal values is expensive to parse and to execute regardless. ## The terms lookup form Instead of an array, the field takes an object: ```json { "terms": { "author_id": { "index": "user_follows", "id": "u-42", "path": "following" } } } ``` Elasticsearch performs an internal get of document `u-42` from `user_follows`, reads the array at `following` from its `_source`, and uses those values as the terms. An optional `routing` parameter tells it which shard to fetch from when the lookup index uses custom routing. Two details matter. First, the values come from `_source`, so the lookup field does not need to be indexed — but `_source` must be enabled on that index. Second, the path may point into nested objects with dotted notation, and the value there must be an array (or a single value) of the terms you want. ## When it is the right call The test is: *does this list already exist as data, and is it big or hot enough that shipping it on every request is wasteful?* Good fits are permission and personalization lists — the categories a tenant may see, the authors a user follows, the product ids in a saved basket, a blocklist. These change independently of the search, they can be large, and they are already stored per user. Keeping them server-side means the application layer does not need to read them before every search, request bodies stay small, and the list has exactly one writer. Poor fits are small, request-specific lists — three status values, a handful of ids the caller already has in hand. There, `terms_lookup` buys nothing and costs an extra fetch. ## The operational costs Every search that uses the lookup does an internal get first. That get is cheap but it is not free, and it lands on whichever shard holds the lookup document. If a thousand searches a second all read the same lookup document, that shard becomes a hotspot. The conventional mitigation is to make the lookup index tiny and ubiquitous: one primary shard, and `auto_expand_replicas` set so a copy lives on every data node. Then the get is always node-local and no single shard carries the whole read load. The second cost is freshness semantics. The document is read at search time, so updates to the list take effect as soon as they are visible to search — which is bounded by the refresh interval, not by any cache you control. That is usually what you want for permissions, but it means an update to the lookup document is not instantly reflected in searches issued in the same second unless you refresh. The third is size. A lookup that resolves to an enormous array is just as expensive to execute as the same array inlined; moving it server-side removes the transfer cost, not the matching cost. ## Alternatives worth naming When the list is genuinely huge and stable, denormalizing is often better: write the group membership onto the searchable documents themselves and filter on a single term. That trades index-time work and reindexing on membership change for a much cheaper query. When the list is derived from another index by a join-like relationship, the honest answer in Elasticsearch is usually denormalization rather than any lookup mechanism, because there is no server-side join. ## Related term-level lookup A close cousin is the `ids` query — `{"ids": {"values": ["1", "2", "3"]}}` — which matches on the `_id` metadata field directly. It is the right tool when you have document ids in hand and want them in one search request alongside other clauses; if you want the documents and nothing else, the multi-get API is a cheaper path than a search.

  • How do you keep a terms lookup index from becoming a hotspot under heavy search traffic?
    Make it a dedicated, tiny index with a single primary shard and `auto_expand_replicas` configured so a replica exists on every data node. The internal get then resolves locally on whichever node runs the search, instead of every search in the cluster hitting one shard. Keeping the lookup documents small keeps that fetch trivial.
  • When is denormalizing the list onto the searchable documents better than terms_lookup?
    When membership changes rarely and the list is large. Writing a group or tenant field onto each document turns the query into a single term filter, which is far cheaper than resolving and matching thousands of terms per search. The price is reindexing affected documents whenever membership changes, so it is a poor fit for volatile lists.
  • What happens if you inline 100,000 values in a terms query?
    The request is rejected once it exceeds `index.max_terms_count`, which defaults to 65,536. Raising the limit is possible but usually the wrong move — parsing and executing that many terms is expensive per shard. Restructure instead: a lookup document, denormalization onto the documents, or a range or prefix predicate that expresses the same set more compactly.

saying these in an interview costs you the question

  • Thinks terms values are analyzed like a match query
  • Believes terms_lookup performs a server-side join across indexes
  • Assumes the fetched list is cached indefinitely
  • Ignores that the lookup shard can become a hotspot
  • Inlines tens of thousands of ids as normal practice

context