Should an Elasticsearch field holding numeric order IDs be mapped as keyword or as a numeric type?
answer
- digits do not imply a numeric type
- which queries will this field receive?
- ranges versus exact lookups
- think about leading zeros
- two different index structures underneath
basics
~10 sMap it as keyword. Numeric types are optimized for range queries, keyword for exact-term lookups, and an identifier is only ever matched exactly. Numeric mapping also destroys leading zeros and invites meaningless arithmetic aggregations.
solid answer
~50 sBeing made of digits is not a reason to use a numeric type. Elasticsearch indexes numeric fields as BKD point trees, which are built for range and sorting; `keyword` fields go into the inverted index, which is built for exact-term lookup. An identifier — order number, customer ID, SKU, postcode, phone number — is queried with `term` and `terms`, never with `>=`, so `keyword` matches the access pattern and generally gives faster term filters. `keyword` also preserves the string exactly, which matters for values with leading zeros or separators that a numeric type silently mangles, and it blocks nonsensical operations like averaging an ID. Use a numeric type when the value is a real quantity you compare, bucket into a histogram, or do arithmetic on. If a field genuinely needs both, index it twice as a multi-field rather than compromising.
code
json · 9 lines{
"properties": {
"order_id": { "type": "keyword" },
"postcode": { "type": "keyword" },
"price_cents":{ "type": "long" },
"client_ip": { "type": "ip" },
"created_at": { "type": "date" }
}
}go deeper
Know the rule of thumb: identifiers you only look up exactly belong in keyword fields, and values you compare, range over or do arithmetic on belong in numeric types.
Explain the structures underneath — inverted index for keyword terms, BKD point trees for numerics — and give the data-fidelity argument about leading zeros and identifier formats that change.
Argue from the actual query mix, size the cost of getting it wrong given that field types are immutable, and offer the multi-field answer when both access patterns are genuinely required.
Set the convention across services: identifiers are strings in the schema everywhere, quantities carry units and the narrowest fitting type, and the mapping review happens before an index template ships rather than after a reindex is needed.
## The question behind the question Interviewers ask this because the naive rule — "it's a number, so use a number type" — is wrong, and the right answer requires knowing that Elasticsearch indexes the two families with completely different data structures aimed at different queries. ## Two structures A `keyword` field goes into the **inverted index**: the value is one term, and the term dictionary maps it straight to a posting list of matching documents. Looking up an exact value is a dictionary seek — precisely the operation the structure exists for. Numeric and `date` fields are indexed as **points in a BKD tree**, a block k-d tree that partitions the value space so a range can be answered by descending to the relevant blocks. That structure is excellent at "all values between X and Y" and at multi-dimensional data, and it is fine but not special at "exactly X", which it has to express as a degenerate range. Both families get `doc_values` for sorting and aggregations, so that is not the deciding factor. ## Match the structure to the access pattern Ask what queries the field actually receives. - **Identifiers** — order number, customer ID, SKU, invoice number, product code — are looked up exactly, sometimes in bulk with `terms`, occasionally by prefix. They are never compared with `>` in a way that means anything: "orders with an ID above 5,000,000" is not a business question. `keyword` is the correct mapping, and Elastic's own guidance is that a numeric-looking field used only for exact matching should be mapped as `keyword`. - **Quantities** — price, latency, byte count, temperature, age — are ranged, bucketed into histograms, summed and averaged. These are numeric types, and the narrowest one that fits the domain (`integer` over `long`, `scaled_float` for fixed-precision money) keeps the index smaller. ## The correctness argument, not just the performance one Performance is the smaller half of the case. Mapping an identifier as numeric changes the data: - **Leading zeros vanish.** A postcode `01234`, an account number `000517`, a country dialling code — as a `long` these become `1234` and `517`. Round-tripping through a numeric type is lossy, and downstream systems comparing against the original string stop matching. - **Separators and mixed forms break ingest.** The day someone's "numeric" ID arrives as `ORD-4471` or `4471-A`, a numeric mapping rejects the document. A `keyword` accepts it. Real identifier schemes change more often than anyone expects. - **Overflow and precision.** Very long digit strings exceed `long`, and floating-point types lose exactness — fatal for an identity comparison. - **Nonsense operations become possible.** With a numeric mapping, nothing stops a dashboard from charting the average customer ID. A `keyword` mapping makes that meaningless query impossible to write. ## When numeric really is right Use a numeric type when the value participates in arithmetic or ordering that means something: a version number you range over, a numeric priority you sort by, a quantity you aggregate. Also consider it when the identifier is genuinely dense and you regularly fetch contiguous blocks of it — a range over a numeric ID is far more efficient than a huge `terms` clause of individual values. And if both patterns are real — exact lookups from the application plus range scans from a batch job — do not compromise. Map the field once with a multi-field so it exists as both a `keyword` and a numeric type, and point each query at the half suited to it. You pay disk for the second copy and buy the right structure for each access pattern. ## Related traps in the same family The same reasoning governs neighbouring choices. An IP address should use the `ip` type rather than `keyword`, because `ip` supports CIDR range matching that a string cannot express. A timestamp should be `date`, not a string, so date math and date-histogram bucketing work. A boolean flag should be `boolean` rather than the strings `"true"`/`"false"`, which cost more and compare inconsistently. In every case the rule is the same: pick the type whose indexed structure answers the queries the field will actually receive, and let the surface appearance of the data — digits, dots, words — be the last consideration rather than the first. ## Deciding in practice Write down the two or three queries the field will serve before choosing. If every one of them is `term`, `terms` or a prefix, map `keyword` and move on. If any of them is a range, histogram or arithmetic aggregation, you need the numeric type — and if some are each, you need both. Because the type cannot be changed on an existing field, this is a five-minute conversation that saves a reindex.
- What if a batch job needs ID ranges but the application needs exact lookups?Index the value twice as a multi-field — a keyword for the term lookups and a numeric sub-field for the ranges — and point each workload at the half that suits it. You pay extra disk for the second set of structures, which is almost always cheaper than forcing one access pattern through the wrong index type.
- Why is a scaled_float often better than a double for prices?scaled_float stores the value as a long multiplied by a fixed scaling factor, so money keeps exact fixed-point precision and compresses far better than a floating-point type. It also sidesteps the classic floating-point comparison surprises when a range boundary sits exactly on a price.
- Should an IP address be mapped as keyword?No — use the ip type. It stores addresses in a form that supports CIDR notation in range and term queries, so you can match a whole subnet in one clause. A keyword mapping can only compare the literal string, which makes any subnet question either impossible or a pile of prefix queries.
saying these in an interview costs you the question
- Says a field of digits must always use a numeric type
- Ignores that a numeric mapping strips leading zeros
- Claims numeric types are always faster for exact matching
- Maps timestamps or IPs as keyword out of habit
- Thinks the field type can be changed later without a reindex