skip to content

Under Elasticsearch dynamic mapping, what type is created for a JSON string like "19.99", and why?

level: middleimportance: should knowfreq 54%

answer

  1. The JSON token decides, not the meaning
  2. Two of the guesses are switchable
  3. One switch is on by default, one is off
  4. Strings produce two mapping entries, not one
  5. Whichever document arrives first sets the type

basics

~20 s

It becomes a text field with a keyword sub-field, because Elasticsearch's numeric_detection defaults to off. Date detection is on by default, so a date-looking string would instead be mapped as date — often the more dangerous surprise.

solid answer

~40 s

Dynamic mapping infers a type from the JSON value, not from what the value looks like it means. `"19.99"` is a JSON string, and because `numeric_detection` is disabled by default, it is mapped as `text` with a `keyword` sub-field (`price.keyword`, with `ignore_above: 256`). The practical consequences are that a `range` query on `price` will not behave numerically and an aggregation must target `price.keyword` and treat values as strings. `date_detection`, by contrast, is **on** by default, so a string matching the recognised date formats becomes a `date` field — which is fine until a later document sends `"N/A"` in that field and the write fails with a parse exception. The type is decided by the **first** document that carries the field, and it cannot be changed afterwards without reindexing.

code

bash · 7 lines
bash
PUT /demo/_doc/1
{ "price": "19.99", "created": "2024-05-01", "qty": 3 }

GET /demo/_mapping
# price   -> text + price.keyword (numeric_detection is off by default)
# created -> date (date_detection is on by default)
# qty     -> long

go deeper

for a junior

Recall that a JSON string becomes text with a keyword sub-field, and that you query the sub-field for exact matches, sorting and aggregations. Knowing the type is fixed once created is enough at this level.

for a middle

Explain the inference table and both detection switches, including which default is on. Be precise that the first document carrying the field decides the type and that changing it later means a reindex.

for a senior

Show how this bites in production: a date-detected field failing writes on a stray value, lexicographic ranges on quoted numbers, and the same field typed differently across indices behind one alias. Name the containment steps.

for a principal

Own the policy that query-relevant fields are declared in templates rather than inferred, and decide where malformed values are handled — rejected at the ingest edge, dropped via ignore_malformed, or allowed to fail writes loudly.

## How the guess is made When `dynamic` is `true` and an unmapped field appears, Elasticsearch picks a type from the JSON token it sees, in roughly this order: - `true`/`false` → `boolean` - a JSON number without a fraction → `long` - a JSON number with a fraction → `float` - an object → `object` - an array → the type of its first non-null element - a string → a `date` if date detection matches, otherwise a numeric type if numeric detection is enabled and matches, otherwise `text` **with a `keyword` sub-field** That last rule is the one to remember: a plain string produces two mapping entries, `field` (analysed, full-text searchable, no doc values) and `field.keyword` (exact-match, sortable, aggregatable, with `ignore_above: 256` so longer values are simply not indexed in the sub-field). ## The two detection switches **`date_detection` defaults to `true`.** Strings matching `dynamic_date_formats` — by default ISO-8601-style values plus a `yyyy/MM/dd` form — are mapped as `date`. This is a convenience that regularly becomes an incident: a free-text field whose first document happens to hold `"2024-05-01"` is locked to `date`, and the next document with `"pending"` in it fails to index. Disabling date detection (`"date_detection": false`) on indices carrying user-supplied strings is a common hardening step. **`numeric_detection` defaults to `false`.** Strings that look numeric stay strings. Turning it on is occasionally useful when an upstream system quotes all values (many CSV and form pipelines do), but it has the same trap in reverse: the field is now `long` or `float`, and a document with `"n/a"` fails. Both are mapping-level properties set beside `properties`, and both only apply to *dynamically* mapped fields. ## First document wins, permanently The critical operational property is that the type is fixed by whichever document first carried the field. In a distributed ingest pipeline you do not control which document that is. Two indices behind the same alias can end up with the same field mapped as `long` in one and `text` in the other, which surfaces later as a cross-index search failure. And because field types are immutable, the fix is a reindex into a corrected mapping, not a mapping edit. A second, subtler consequence: `"19.99"` mapped as `text` sorts and ranges lexicographically through the `keyword` sub-field, so `"9.5"` sorts *after* `"19.99"`. Nothing errors; the numbers are simply wrong. Silent wrongness is the reason interviewers ask this question. ## Arrays and nulls An array takes the type of its first non-null element, so `[null, null, 3]` maps as `long` while `["3", 3]` maps as `text`. A field whose first document has `null` (or `[]`) is not mapped at all — the field is skipped, and the type is decided by the next document that carries a real value. That is why a smoke-test document full of nulls does not "pre-create" a mapping. ## Containing it The options, roughly in order of how much you should prefer them: 1. **Declare the fields explicitly.** For any field an application actually queries, an explicit mapping in an index template is the correct answer; dynamic inference is for the tail, not the core. 2. **Turn off `date_detection`** on indices that receive arbitrary strings. 3. **Use dynamic templates** to impose a house rule — for example, map every dynamically detected string as `keyword` only, which both fixes the type and halves the number of mapping entries. 4. **Set `ignore_malformed`** on numeric and date fields where a bad value should be dropped rather than fail the document. It is a per-field mapping parameter, and it makes the offending field unsearchable for that document while the rest of it still indexes. ## Checking what you got `GET /<index>/_mapping` after the first documents land is a cheap habit, and in CI a test that asserts the resulting mapping catches drift before production does. If you need to know how a value *would* be treated before committing, index into a throwaway index first — dynamic mapping is easy to inspect and impossible to undo in place.

  • Why do teams disable date_detection on indices that carry user-supplied strings?
    Because a single early document whose string happens to look like a date locks the field to `date` for the life of the index. Every later value that is not parseable — an empty string, `"N/A"`, a free-text note — fails the whole document with a parse exception. Disabling detection keeps such fields as text/keyword, which accepts anything.
  • What type does a dynamically mapped field get when its first document has the value null?
    None — the field is skipped entirely and does not appear in the mapping. `null` and an empty array carry no type information. The mapping is created by the first document that supplies a real value, which is why null-heavy smoke-test documents never pre-establish a schema.

saying these in an interview costs you the question

  • Assumes a quoted number is auto-detected as numeric
  • Thinks numeric_detection is on by default
  • Believes a wrong dynamic type can be edited in place
  • Forgets a dynamic string also gets a keyword sub-field
  • Says sorting the text field's keyword sub-field orders numerically

context