skip to content

What does Solr's schemaless mode do when it indexes a document containing an undeclared field?

level: middleimportance: should knowfreq 54%

answer

  1. it is two features stacked, not one mode
  2. something writes the schema at index time
  3. which document decides the type
  4. the first value seen fixes the type forever

basics

~20 s

It guesses a type from the value in the first document that carries the field, then permanently adds an explicit field declaration to the managed schema through an update processor chain. The guess is never revisited for later documents.

solid answer

~40 s

"Schemaless" is not a server mode but two features stacked. First, a **managed schema** (`ManagedIndexSchemaFactory`, mutable) makes the schema writable at runtime through the Schema API. Second, an update request processor chain — `add-unknown-fields-to-the-schema` in the `_default` configset — sits in front of indexing: it parses values with `ParseBooleanFieldUpdateProcessorFactory`, `ParseLongFieldUpdateProcessorFactory`, `ParseDoubleFieldUpdateProcessorFactory` and `ParseDateFieldUpdateProcessorFactory`, and then `AddSchemaFieldsUpdateProcessorFactory` maps the resulting Java type to a Solr fieldType and issues an add-field. From then on the field is an ordinary explicit declaration. The danger is that the guess comes from **one** document: an id like `"00123"` may become a long, a version string may become a date, and the next document that violates the guess fails to index. Production practice is to define fields deliberately via the Schema API and switch guessing off with the `update.autoCreateFields` property.

code

bash · 11 lines
bash
# turn field guessing off while keeping the managed schema editable
curl http://localhost:8983/solr/mycoll/config -H 'Content-type:application/json' \
  -d '{"set-user-property":{"update.autoCreateFields":"false"}}'

# declare the fields deliberately instead
curl http://localhost:8983/solr/mycoll/schema -H 'Content-type:application/json' -d '{
  "add-field": [
    {"name":"price", "type":"pfloat", "stored":true, "docValues":true},
    {"name":"sku",   "type":"string", "stored":true, "docValues":true}
  ]
}'

go deeper

for a junior

Know that Solr can create fields for you when it meets unknown ones, and that this is a convenience for getting started rather than how production collections are run.

for a middle

Explain the two pieces — a mutable managed schema plus an update processor chain that parses values and calls add-field — and name at least one way a guess goes wrong, such as a zero-padded id becoming a number.

for a senior

Be able to argue for switching guessing off, show the Schema API calls or configset you would ship instead, and describe the reindex you owe when a wrong type has already been baked in. Field explosion from a misbehaving producer should be on your risk list.

for a principal

Treat the schema as a versioned interface between producers and search. Decide whether types are declared in a reviewed configset, whether dynamic field suffixes are the contract, and how a schema change rolls out alongside the reindex it forces across environments.

## Schemaless is a stack, not a mode Nothing in Solr is labelled "schemaless engine". What the `_default` configset ships is a combination of two independent capabilities, and understanding the split is what the question is really testing. **1. A managed schema.** `solrconfig.xml` declares `<schemaFactory class="ManagedIndexSchemaFactory"/>`, which makes the schema file (`managed-schema.xml` in Solr 9; `managed-schema` with no extension in 8.x) writable at runtime and exposes the Schema API at `/solr/<collection>/schema`. The alternative, `ClassicIndexSchemaFactory` reading a hand-edited `schema.xml`, makes the Schema API read-only. In SolrCloud the managed schema lives in the configset in ZooKeeper, and a successful change triggers a reload of the collection's cores. **2. A field-guessing update chain.** An update request processor chain runs before documents reach the index writer. The chain named `add-unknown-fields-to-the-schema` typically contains, in order: a UUID processor for missing ids, `RemoveBlankFieldUpdateProcessorFactory`, `FieldNameMutatingUpdateProcessorFactory` (which sanitizes illegal characters in field names), the parse processors for booleans, longs, doubles and dates, and finally `AddSchemaFieldsUpdateProcessorFactory`. ## What actually happens on an unknown field The parse processors try to convert the incoming string into a richer Java type; `"42"` becomes a Long, `"2026-01-01T00:00:00Z"` becomes a Date, `"true"` becomes a Boolean, and anything else stays a String. `AddSchemaFieldsUpdateProcessorFactory` then consults its `typeMapping` list, which maps each Java class to a Solr fieldType, and adds that field to the managed schema. The schema reloads and the document is indexed against a now-explicit field. Crucially, this happens **once**, driven by whichever document first carries the field. Nothing re-examines the decision afterwards. ## The failure modes - **A narrow first value.** The first document has `"quantity": 3`, so the field becomes a long. A later document sends `3.5` and fails with a number-format error — the whole update request may be rejected. - **Accidental coercion.** Postal codes, phone numbers and zero-padded ids like `"00123"` parse as numbers, losing leading zeros and breaking exact lookups. Version strings and free text that looks date-like can become dates. - **Text where you wanted a keyword.** Plain strings land in an analyzed text type in the stock configuration, so exact matching, sorting and faceting on them behave unexpectedly unless the configuration also copies into a string field. - **Field explosion.** A malformed producer that emits `attr_<uuid>` names will happily create thousands of real fields, each with its own term dictionary and possibly docValues, bloating the index and the heap. - **Irreversibility.** Removing a wrong field or changing its type does not fix the documents already indexed under it. You reindex. - **Concurrency.** Under parallel indexing, several requests can race to add the same field; Solr resolves the conflict by retrying, but the retries add latency during bulk loads. ## Turning it off In the `_default` configset the chain is gated by the `update.autoCreateFields` user property, which you flip through the Config API: ```bash curl http://localhost:8983/solr/mycoll/config -H 'Content-type:application/json' \ -d '{"set-user-property":{"update.autoCreateFields":"false"}}' ``` The managed schema remains mutable, so you keep the Schema API and lose only the guessing. ## The production pattern Use guessing as a discovery tool: point it at a sample of real data in a throwaway collection, look at what schema it produced, then hand-write the fields you actually want. Ship those either as a configset containing a prepared `managed-schema.xml` or as a sequence of Schema API calls in your provisioning script: ```json {"add-field":{"name":"price","type":"pfloat","stored":true,"docValues":true}} ``` That gives you deliberate types, code-reviewed schema changes, and a loud rejection when a producer sends something unexpected — which is what you want in production, where a silent type guess is a bug that surfaces weeks later as "search stopped finding these documents". Dynamic fields are often the right middle ground: the schema stays fixed and reviewable, while suffix conventions like `*_i` and `*_txt` let producers introduce new attributes without a schema change and without a guess.

  • A field was guessed as a long and a later document sends 3.5 for it. What happens?
    That document fails to index with a number-format error, and depending on how the batch was sent the whole update request can be rejected. Solr does not widen an existing field to a broader type; the schema entry is already fixed. Fixing it means changing the field type and reindexing everything, because documents already stored under the old type keep the terms they were indexed with.
  • How is a managed schema different from a classic schema.xml?
    A managed schema is declared with `ManagedIndexSchemaFactory` and is writable at runtime via the Schema API, with changes persisted to the schema file — in SolrCloud, into the configset in ZooKeeper, followed by a collection reload. `ClassicIndexSchemaFactory` reads a hand-edited `schema.xml` and makes the Schema API read-only. Field guessing needs the managed, mutable form; the reverse is not true, since you can run a managed schema with guessing disabled.
  • Why do teams prefer dynamic fields over schemaless guessing for open-ended attributes?
    Because the type is chosen by the field name the producer picks, not inferred from one sample value. A field named `weight_d` is unambiguously a double whatever the first document contains, the schema stays fixed and reviewable, and an unexpected name either matches a pattern or fails loudly. Guessing, by contrast, silently freezes a decision made from a single document.

saying these in an interview costs you the question

  • Says Solr infers the type per document, not once
  • Thinks the guessed type widens automatically for later values
  • Confuses schemaless guessing with dynamic field patterns
  • Believes changing a guessed field's type fixes existing documents
  • Assumes a managed schema implies field guessing is on

context